Senior Site Reliability Engineer
Job description
About the role
Join the Future of Supply Chain Intelligence
Powered by Agentic AI. At Resilinc, we engineer autonomous systems that revolutionize supply chain risk management. Our agentic AI empowers global enterprises to predict disruptions, assess impact, and execute real-time responses before operations suffer. Recognized as a 2025 Gartner® Magic Quadrant™ Leader, we protect mission-critical operations for leaders in life sciences & pharma, aerospace & defense, high tech, and automotive. Join a team redefining global resilience through high-impact work. Our strength lies in our people. We operate as a fully remote, mission-led team ensuring life-saving products and critical goods reach their destinations without delay. We cultivate a collaborative, empowering culture where you can be an agent of change. Tackle critical global challenges through work that matters and shapes the future of supply chain intelligence. Explore our blog to learn how we are transforming the world's most critical supply chains: *Global Supply Chain Risks 2026: Act Faster | TEC*. Resilinc | Innovation with Purpose. Intelligence with Impact.
We are seeking an experienced Site Reliability Engineer (SRE) to ensure the reliability of our platform at scale. This India-based, night-time role provides dedicated coverage aligned with US business hours, closing a critical support gap. You will own production incidents, enhance system reliability, reduce MTTR, and drive automation across a modern cloud-native stack. This position offers direct visibility into business-critical operations and impactful work on large-scale distributed systems.
Key Responsibilities
You will design control flows that maintain global supply chain visibility during incidents, ensuring systems remain online when it matters most. Teams will coordinate response patterns so life-saving cargo never stalls within silent systems. Your work will translate risk signals into immediate action before partners experience disruption. You will route data through pipelines while confirming integrity for critical modules. Designing deployment sequences will require strict adherence to change windows and robust rollback paths. Automating alert responses will enable on-call staff to focus on complex investigations. Observing queue depths and latency will be essential to maintaining service stability under peak load. Documenting runbooks will clarify ownership during multi-team incident handling. You will test recovery drills to validate backups and restore procedures for essential stores. Coordination with partners will align monitoring standards across shared supply chain layers. Reviewing logs and traces will close gaps affecting uptime or data correctness.
Requirements
You must hold a minimum of 3 years of experience managing cloud infrastructure in production environments. Hands-on work with Azure services and Kubernetes orchestration in live settings is required. You must understand Kafka messaging patterns and their role in moving events between systems. The role requires night-shift availability to align with United States business hours. Writing automation scripts using tools that interact with modern stacks is essential.
Nice to have
Experience with Resilinc or supply chain platforms is beneficial but not mandatory.
What you'll do
- Meet the bar Confirm that the role is based in India and operates on a night-time schedule aligned with US business hours.
Verify that the position is full-time and that the candidate will own production incident response and system reliability improvements.
Check that the listed tools and skills include Azure, Kubernetes, and Kafka as core requirements for the role.
Review the responsibilities to ensure they cover designing control flows, coordinating response patterns, and translating risk signals into action.
Ensure that the requirements specify a minimum of 3 years of cloud infrastructure management and hands-on Kubernetes experience.
Confirm that the role demands night-shift availability and script-writing automation skills for modern stacks.
Validate that the nice to have section mentions prior experience with Resilinc or supply chain platforms as beneficial but not mandatory.
Read the practical notes to understand that this listing is specific to India and that official apply page details must be consulted for policy and compensation.
Follow the note that directs interested candidates to the official apply page for complete information on duties, pay, and location.
Ensure that repeated calls to apply on the official posting are acknowledged to confirm current details before proceeding.
Understand that this is a mission-led, fully remote role operating within a collaborative and empowering culture focused on global supply chain resilience.
Recognize that the position provides visibility into business-critical operations and involves work on large-scale distributed systems in a cloud-native environment.
Acknowledge that success in this role requires reducing MTTR, driving automation, and maintaining service stability under peak load conditions.
Realize that documenting runbooks, testing recovery drills, and coordinating with partners are essential day-to-day activities for this engineer.
Accept that the candidate must meet the bar outlined Confirm that the role contributes to life-saving cargo visibility and that the work has direct impact on critical global supply chain operations.
Understand that the candidate will be part of a team redefining supply chain intelligence through agentic AI and autonomous systems.
Know that Resilinc is recognized as a 2025 Gartner® Magic Quadrant™ Leader protecting mission-critical operations across several major industries.
See that the position closes a critical support gap by providing dedicated night-time coverage aligned with US business hours from an India-based location.
Value the requirement for hands-on experience with Azure services, Kubernetes orchestration, and Kafka messaging patterns in live production settings.
Respect the expectation that the engineer will own production incidents, enhance reliability, and automate alert responses to improve system performance.
Note that the role involves designing deployment sequences with strict change windows and robust rollback paths to ensure continuity.
Appreciate that the work includes translating complex risk signals into immediate action to prevent disruptions for supply chain partners.
Understand that the position is full-time and based in India, with compensation details to be confirmed See that the role emphasizes a mission-led culture focused on ensuring life-saving products and critical goods reach destinations without delay.
Recognize that the successful candidate will contribute to a collaborative, empowering culture as an agent of change within a globally distributed team.
Acknowledge that the position offers direct visibility into business-critical operations and impactful work on large-scale distributed systems.
Understand that the role supports the company's vision of engineering autonomous systems that revolutionize supply chain risk management.
Know that the position requires the ability to work nights to align with United States business hours, ensuring continuous system reliability.
Confirm that the tools and skills section lists Azure, Kubernetes, and Kafka as essential technologies for the role.
Realize that the responsibilities include documenting runbooks, testing recovery drills, and coordinating monitoring standards with partners.
Understand that the role demands script-writing abilities and experience with modern stacks to automate alert responses effectively.
See that the requirements include a minimum of 3 years managing cloud infrastructure and hands-on work with live Kubernetes environments.
Recognize that the nice to have qualification mentions experience with Resilinc or supply chain platforms as an advantage.
Follow the practical notes that direct candidates to Understand that the official apply page is the source for verifying remote policy, pay, and start date specifics.
Accept that repeated applications to the official posting may be necessary to confirm the most current information.
Value the note that this Senior Site Reliability Engineer listing specifically names India as the location.
See that the role is part of a mission to transform supply chain intelligence through agentic AI and autonomous systems.
Recognize that the successful candidate will own production incidents and drive automation to reduce MTTR.
Understand that the position is full-time and operates within a night-time schedule to support US business hours.
Confirm that the role requires designing control flows to maintain visibility during incidents and ensuring system uptime.
Acknowledge that the work involves routing data through pipelines and confirming integrity for critical modules.
See that the role includes designing deployment sequences with adherence to change windows and rollback paths.
Understand that the responsibilities involve automating alert responses to enable on-call staff to focus on complex investigations.
Recognize that the position requires observing queue depths and latency to maintain service stability under peak load.
Value the requirement to document runbooks for clarity during multi-team incident handling and testing recovery drills for essential stores.
Understand that coordination with partners is necessary to align monitoring standards across shared supply chain layers.
Accept that reviewing logs and traces is part of the role to close gaps affecting uptime or data correctness.
Know that the position is located in India and is a night-time role aligned with US business hours.
Confirm that the role is full-time and involves ownership of production incidents, system reliability, and automation efforts.
Understand that the tools and skills section specifies Azure, Kubernetes, and Kafka as core technologies for the position.
See that the requirements include a minimum of 3 years of cloud infrastructure management and hands-on Kubernetes experience in production.
Recognize that the role demands night-shift availability and script-writing skills for modern stacks.
Acknowledge that the nice to have section mentions experience with Resilinc or supply chain platforms as beneficial.
Follow the practical notes that direct candidates to the official apply page for complete details on policy, pay, and location.
Understand that the official apply page is the definitive source for confirming duties, compensation, and location specifics.
Accept that repeated applications to the official posting may be needed to verify current information.
Value the mission-led culture focused on ensuring life-saving products and critical goods reach destinations without delay.
Recognize that the role provides visibility into business-critical operations and involves impactful work on large-scale distributed systems.
Understand that success in this role requires reducing MTTR, driving automation, and maintaining service stability under peak load.
See that the responsibilities include documenting runbooks, testing recovery drills, and coordinating monitoring standards with partners.
Acknowledge that the role demands script-writing abilities and experience with modern stacks to automate alert responses effectively.
Confirm that the position is full-time and operates within a night-time schedule to support US business hours from India.
Recognize that the tools and skills section lists Azure, Kubernetes, and Kafka as essential technologies for the role.
Understand that the requirements include a minimum of 3 years managing cloud infrastructure and hands-on work with live Kubernetes environments.
See that the nice to have qualification mentions expe