Senior Software Engineer, Cloud Development
Job description
Senior Software Engineer, Cloud Development
About the role
Mozilla is seeking a Senior Software Engineer to join our AI Platform team. This role involves building the core infrastructure that supports intelligent features across Mozilla's products. You will work on systems for training models, high-volume inference, GPU management, and secure, privacy-focused AI operations at a global scale.
Key facts
What you'll do
Design, construct, and maintain essential platform services and APIs for deploying and serving production workloads.
Take full responsibility for service reliability, driving enhancements in availability, scalability, performance, and operational efficiency.
Lead initiatives to optimize backend services for throughput, latency, and cost-effectiveness across distributed infrastructure.
Develop and manage Kubernetes-based workloads, including GitOps deployment processes, environment configuration, and resource usage optimization.
Oversee and enhance critical aspects of the service lifecycle, such as packaging, versioning, testing methodologies, validation, and automated deployments.
Implement and refine observability practices, including metrics, logging, tracing, and alerting, to improve the visibility and resilience of backend services and pipelines.
Collaborate with product, infrastructure, and data teams to create scalable platform capabilities that enable new product features.
Contribute to technical design discussions, suggest architectural improvements, and mentor less experienced engineers through code reviews and knowledge sharing.
Participate in and help improve operational procedures, including incident response, on-call duties, and post-incident analysis.
Requirements
Bachelor's degree with 4-6 years of relevant professional experience, or a Master's degree with substantial practical experience building and operating production systems, or equivalent work experience.
Proficiency in modern Python, with a track record of writing clean, maintainable code and utilizing a fast toolchain (dependency management, linting, formatting, type checking, pre-commit hooks) for building libraries and CLIs that produce structured data.
Advanced experience with database deployment and administration, with a preference for familiarity with Postgres.
Demonstrated experience deploying and operating workloads in cloud environments, including production-grade infrastructure on GCP and GKE (artifact registries, managed caches, networking, internal load balancing, VPC, DNS, and separation of non-production and production environments).
Practical experience with Kubernetes and Helm, including writing charts for multi-environment deployment with per-environment configuration and progressive feature rollouts.
Experience using Terraform for infrastructure provisioning across environments, including schema validation and plan review via pull requests.
Experience designing and operating scalable APIs that perform well under load, including health and readiness checks, authentication, and graceful startup and shutdown procedures.
Experience with Grafana or similar tools for metrics, dashboards, and analyzing application and infrastructure health during deployments.
Strong analytical and problem-solving abilities, with a capacity to debug performance and reliability issues in distributed systems.
Effective communication skills, with experience collaborating across engineering, product, and infrastructure teams.
Experience with on-call responsibilities, including participation in incident response and post-incident reviews.
Nice to have
Experience with Ray or Ray Serve for GPU-accelerated model serving, including setting resource requests and replica counts based on available hardware.
Experience building stateless machine learning services like embedding or similarity models, including multi-model loading, runtime device selection, batch APIs, and managing trade-offs between model caching and cold starts.
Experience operating a multi-provider LLM gateway, including routing between providers, migrating models, and combining self-hosted and third-party serving.
Familiarity with containerization and orchestration systems in production beyond core Kubernetes/Helm usage.
Exposure to privacy-preserving machine learning techniques, security best practices, or responsible AI system design principles.
Contributions to open-source infrastructure projects or leadership in developing reusable internal tools.
Skills & tools
Python
GCP
GKE
Kubernetes
Helm
Terraform
Grafana
Postgres
Ray (Ray Serve)
Practical notes
Mozilla offers generous performance-based bonuses, comprehensive medical, dental, and vision coverage, significant retirement contributions with immediate vesting, quarterly wellness days, country-specific holidays plus a birthday off, a home office stipend, an annual professional development budget, a quarterly well-being stipend, substantial paid parental leave, and an employee referral bonus program. Other benefits vary by country.