
AI Applications Ops Lead, GPS
Job description
About the role
The is responsible for architecting the production lifecycle of AI systems that serve national priorities across multiple sovereign territories. This position owns the design, implementation, and long-term reliability of full-stack AI deployments that directly support government agencies on the international stage. The role requires ownership over the stability and security of model serving infrastructure, ensuring that AI capabilities remain performant and available under strict public sector demands. You will manage the intricate connectivity between user interfaces, backend APIs, and the core AI execution engines to maintain seamless operations. A core part of this position involves developing and enforcing automated monitoring strategies that track model performance and data drift across distributed global sites. You will act as the primary incident commander for production issues, driving rapid resolution and establishing preventative guardrails to avoid future disruptions. This role will communicate highly technical performance metrics and system health to senior government officials, translating complexity into clear operational narratives. You will collaborate closely with engineering and machine learning teams to translate field observations into concrete improvements for future technical architecture and strategic roadmaps.
Key facts
What you'll do
- Assume ownership for the long-term reliability and performance of AI deployments within government environments, ensuring strict adherence to service level objectives.
- Oversee the end-to-end health of production platforms, maintaining robust connectivity between user interfaces, APIs, and AI cores to prevent service fragmentation.
- Design and implement automated monitoring frameworks that provide real-time visibility into model performance and data drift across international regulatory jurisdictions.
- Navigate and uphold compliance standards within various global regulatory and compliance frameworks, ensuring all operations meet legal and security requirements.
- Serve as the designated incident commander for critical production issues, executing rapid response actions and implementing strategic preventative guardrails.
- Translate intricate technical performance data into accessible narratives for senior government officials, ensuring stakeholders understand risk, status, and mitigation strategies.
- Partner with engineering and machine learning teams to analyze field deployment data, driving informed decisions that shape future technical architecture.
- Lead the management of the complete request lifecycle, optimizing system components to handle government scale without degradation in performance.
- Establish and maintain operational runbooks that standardize responses to anomalies, ensuring consistency across geographically dispersed support teams.
- Champion the adoption of modern infrastructure patterns that enhance the resilience and observability of AI applications in sensitive public sector contexts.
Requirements
- Possess 6+ years of professional experience in Site Reliability Engineering, Full-Stack Development, or MLOps roles, demonstrating a track record of managing complex systems.
- Bring a documented background working within the public sector, understanding the unique constraints and priorities of government technology initiatives.
- Show direct experience with sovereign AI deployment strategies and adherence to international government security standards and data governance policies.
- Demonstrate proficiency in maintaining production-grade systems, with a comprehensive understanding of the entire request lifecycle from ingress to egress.
- Exhibit the ability to dissect complex system performance degradation and articulate the root causes and implications to non-technical stakeholders effectively.
- Hold a proven capacity to operate comfortably in ambiguous, high-stakes environments where regulatory compliance is a primary driver of technical decisions.
- Have hands-on experience with infrastructure that supports global deployment, including considerations for latency, data residency, and jurisdictional compliance.
- Display strong judgment in prioritizing operational tasks to ensure continuity of service for critical government applications.
Nice to have
- Deep technical expertise in modern AI infrastructure, including the intricacies of model serving, scaling, and optimization in distributed environments.
Skills & tools
- Kubernetes
- Vector databases
- Agentic development
- LLM observability tools
Practical notes
For Doha-based roles: Candidates must provide personal data for visa and residency processing as required by Qatari authorities. Scale AI supports the application process, though final issuance remains at the discretion of the government.
There is a mandatory 90-day waiting period for candidates reapplying for the same position.
Reasonable accommodations are available for applicants with disabilities by contacting accommodations@scale.com.