Jobshare - Sr Lead Software Engineer
Job description
About the role
This is an exciting opportunity where we're exploring local Sr Lead Security Engineering talent who need flexibility, with the potential to present qualified candidates to managers for part-time or job-share roles. Part-time hours where two people share the responsibilities. We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. You will own the design and execution of a next-generation reliability and security operating model for a critical global business line. The role requires deep technical leadership in defining how services are observed, protected, and run at scale. You will partner closely with production support and development teams to instill reliability thinking across the entire service lifecycle. Success in this position will be measured by your ability to reduce operational toil while improving system resilience and user experience. You will leverage modern practices and tooling to create a robust, scalable, and auditable approach to managing complex production systems.
Key facts
What you'll do
- Define the SRE vision, north-star outcomes, and multi-year roadmap for the Production Management team, aligned to both CIB and JPM Global Technology priorities.
- Establish the SRE operating model across global regions, including ways of working, intake processes, prioritization frameworks, and engagement with engineering teams and production support.
- Partner with business-aligned Production Support leads to embed SRE practices consistently and act as a force multiplier by coaching them on reliability thinking, prioritization, and engineering out operational load.
- Build and develop a small, high-impact core SRE team and foster a virtual SRE community of practice that scales reliability improvements across many application flows and services.
- Define and implement standards for service cataloging, SLO/SLI frameworks and error budgets, incident response maturity, blameless post-incident reviews, resiliency patterns, capacity, performance, and scalability engineering.
- Drive service reviews with evidence-based reporting on availability, latency, incident trends, MTTR and MTTD, change failure rate, and customer impact to guide improvement initiatives.
- Champion AI adoption and deliver AI-enabled capabilities that reduce operational toil and improve the speed and quality of response for reliability and security workflows.
- Set direction for observability across logs, metrics, and traces, including instrumentation standards, golden signals, and end-user journey monitoring to improve alert quality and routing.
- Build strong partnerships with application development teams, platform and infrastructure partners, and governance functions to ensure cohesive delivery of reliability outcomes.
- Communicate clearly and credibly at all levels, translating technical concepts into actionable insights for engineers, senior technology leaders, and business stakeholders.
- Use enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis while validating outputs and handling operational data according to sensitivity and security requirements.
- Lead reuse-first adoption of AI-assisted reliability workflows across the software development lifecycle and toolchain, including CI/CD quality checks, test and validation automation, and operational readiness.
- Ensure traceability, auditability, resiliency, and security controls are embedded into all reliability initiatives and process improvements.
- Continuously evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align with resiliency and security expectations.
- Serve as a technical authority and thought leader, mentoring peers and junior engineers on best practices in SRE, security, and reliability engineering.
Requirements
- 10+ years of experience in technology support, production/application support, DevOps, or infrastructure management.
- Demonstrated experience leading SRE/reliability engineering or production engineering transformations in a complex enterprise environment.
- Strong engineering background with the ability to design, build, and deliver automation and reliability solutions that scale.
- Fluency and expertise in Python for developing automation, tooling, and analysis scripts.
- Deep practical knowledge of SLOs and SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces, and distributed systems troubleshooting.
- Proficiency and experience with telemetry collection using tools and standards such as Prometheus, Open Telemetry, Datadog, Dynatrace, and Splunk.
- Experience delivering automation at scale through scripting, workflow automation, runbook automation, and CI/CD-integrated guardrails.
- Proven leadership skills, including influencing without authority, coaching leaders, and building communities of practice within and across technology organizations.
- Strong judgment around risk, security, and controls, particularly when applying AI to production workflows and sensitive operational data.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows, such as incident investigation support and knowledge capture.
- A track record of strong validation habits and awareness of data sensitivity, confidentiality, and regulatory requirements.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align with resiliency and security expectations.
- Excellent written and verbal communication skills to articulate strategy, rationale, and technical decisions to both technical and non-technical audiences.
- Willingness to collaborate in a part-time, job-share arrangement where responsibilities are divided effectively with another professional.
Nice to have
- Familiarity with Ath
Practical notes
- Engagement is part_time with shared responsibilities in a job-share arrangement.
- No specific hours, travel, visa, or deadline information is provided in this description.