Senior Systems Support Engineer
Job description
About the role
You will own the full lifecycle of complex production incidents across multi-tier client environments, coordinating with development and operations teams to restore service swiftly and thoroughly. You will apply deep expertise in application monitoring and observability tooling to analyze metrics, generate actionable reports, and drive corrective actions that reduce future risk. You will navigate intricate application architectures to isolate root causes for business-impacting issues while adhering to strict standards that enhance operational efficiency, stability, and availability. You will leverage continuous delivery practices to deploy and evolve high-quality software as early as possible, collaborating in value-driven agile teams that build innovative client experiences. You will utilize diverse logging techniques and levels to power alerting, monitoring, and forensic analysis during incident investigations. You will employ DevOps tools and practices to automate, manage, and run software deployments reliably and at scale. You will mentor less experienced peers by sharing technical knowledge and leadership behaviors that strengthen the overall support capability. You will apply the latest insights from the Thoughtworks Technology Radar to evaluate, adopt, and adapt emerging solutions for client problems.
Key facts
What you'll do
You will use your skills in incident management processes and tools, application monitoring metrics and tooling to generate reports and take corrective actions. You will understand complex application systems and find your way through them to debug a business impacting issue. You will follow standards and best practices to bring operational efficiencies, stability and availability of the system. You will use continuous delivery practices to evolve, support and deliver high-quality software, as well as value to end customers, as early as possible while working in a collaborative, value-driven teams to build innovative customer experiences for our clients. You will leverage your knowledge regarding the different logging techniques (various levels) and use them for alerting, monitoring and identifying the root cause of incidents. You will efficiently use DevOps tools and practices to deploy and run software. You will act as a mentor for less experienced peers through both your technical knowledge and leadership skills. You will apply the latest technology thinking from our Technology Radar to solve client problems. You will partner with development teams to design, implement, and validate resilient deployment pipelines that support rapid and safe delivery of features. You will conduct post-incident reviews to extract learnings and drive improvements in processes and tooling. You will collaborate with operations and site reliability teams to align on monitoring, alerting, and capacity strategies. You will translate business requirements into technical support plans that ensure service continuity and performance targets. You will maintain and evolve runbooks and operational documentation to reflect current system behavior and response procedures. You will participate in on-call rotations as needed to provide 24x7 coverage for critical production systems. You will contribute to the professional growth of team members by sharing hands-on troubleshooting techniques and decision-making frameworks.
Requirements
You have experience working on programming languages such as Java or .Net, an understanding of cloud platforms such as AWS, Azure or GCP and scripting languages such as Python or Powershell. You have a high-level understanding of various architectures such as monolithic, N-tier, layered, microservices and serverless. You must possess strong debugging and triaging skills to troubleshoot code effectively. You have experience working with relational databases such as MS SQL, MySQL or PostgreSQL. You have experience working with containerization tools such as Docker and CI/CD tools such as Jenkins and Azure pipelines. You have experience working with relational or non-relational databases. You have an understanding of application monitoring tools such as DataDog, Prometheus or Grafana, the different metrics that go with it and are able to generate reports and take corrective actions. You have the ability to ensure that the deliverables, namely bug fixes and enhancements to the existing codebase, are high-quality and well-tested. You are comfortable with Agile methods, such as Scrum and/or Kanban. You are willing to be part of a rotation- and need-based 24x7 available team. You can navigate ambiguous requirements and make sound decisions with limited information while maintaining a methodical approach to problem solving. You communicate clearly and can articulate technical concepts to both technical and non-technical stakeholders. You take ownership of your work and demonstrate accountability for outcomes across distributed systems and teams. You are proactive in identifying risks and proposing mitigations before issues escalate into major incidents.