Software Development Engineer II, Cloud Platform
Job description
About the role
Mapbox is seeking a Software Development Engineer II to join the Cloud Platform team in shaping the infrastructure that powers our real-time location platform. In this role, you will own the design, operation, and evolution of cloud-native platforms and deployment workflows. You will work alongside a globally distributed team to build tools that enable hundreds of engineers to deliver reliable and secure services. This position requires thoughtful integration of AI into engineering workflows and infrastructure decision-making. You will be responsible for driving operational excellence through automation, observability, and robust deployment practices. Your work will directly impact the scalability, security, and developer experience for Mapbox services used by millions of developers worldwide.
Key facts
What you'll do
- Actively onboard AWS resources to the declarative gitops-based framework utilizing Terraform and Terragrunt.
- Maintain and troubleshoot legacy cloud infrastructure in AWS that is deployed with Cloudformation/CDK and utilizes ECS, Lambda, EMR, and related services.
- Architect and promote Kubernetes deployments for new services across diverse workload patterns.
- Lead migration of deployment pipelines from ECS and Cloudformation to EKS and ArgoCD to improve reliability and developer velocity.
- Architect a centralized CI pipelines framework utilizing GitHub Actions and Runs-on to standardize builds and deployments.
- Broadly influence and lead the Mapbox Cloud Platform strategy around AWS architecture, open-source tools, and internal frameworks.
- Configure and maintain a comprehensive observability platform, such as Datadog or Observe, to enable real-time monitoring, alerting, and analytics for production services.
- Promote a culture of operational excellence by testing and monitoring our systems and code, and providing on-call support for the platform services in production.
- Document your work and decision-making processes, and lead presentations and discussions in a way that is easy for others to understand and adopt.
- Uphold a culture of collaboration, transparency, creativity, inclusion, and data-driven decisions within cross-functional teams.
- Partner with application teams to ensure platform capabilities meet evolving business requirements while maintaining security and compliance standards.
- Identify opportunities for automation, cost optimization, and performance improvements across the cloud infrastructure stack.
- Mentor engineers on best practices for infrastructure-as-code, container orchestration, and cloud operations.
- Participate in incident response activities, leading post-incident reviews to drive continuous improvement of system resilience.
Requirements
- 5+ years experience leveraging infrastructure-as-code frameworks to manage AWS infrastructure using Terraform, Terragrunt, Atlantis, CDK.
- 4+ years experience orchestrating containerized workloads at scale using EKS, ECS, and related Kubernetes ecosystems.
- 4+ years experience managing scalable CI/CD frameworks in a distributed engineering organization using GitHub Actions and related pipeline technologies.
- Strong expertise with Kubernetes, ArgoCD, Istio, and service mesh concepts for production-grade deployments.
- Proven ability to design and develop cost efficient, secure, and durable solutions on AWS using EKS, ECS, EC2, Lambda, Fargate, CloudFront, IAM, Route53, DynamoDB, and other core services.
- Proficient in at least one programming language, such as Python, Nodejs, or GoLang, to develop tooling and automate infrastructure operations.
- Experience configuring and managing observability systems in a distributed large-scale environment using Datadog, CloudWatch, or similar monitoring platforms.
- Experience with incident response practices including blameless post-mortems, root cause analysis, and resilience engineering concepts.
- A desire to share your expertise through documentation, mentorship, and both written and vocal communication with engineering teams.
- Ability to work asynchronously and independently with minimal supervision, lead by example, and make decisions based on priorities and business goals.
- Strong understanding of security and compliance principles as they apply to cloud infrastructure and data protection.
- Willingness to engage in continuous learning to keep up with rapidly evolving cloud technologies and industry best practices.
- Commitment to following Mapbox engineering processes, tools, and standards to ensure consistency and quality across teams.
- Readiness to collaborate effectively with team members across different time zones and cultural backgrounds in a globally distributed environment.
- Passion for building and maintaining reliable platform services that enable customer innovation and business growth.
Nice to have
- Experience contributing to open-source projects relevant to cloud infrastructure, CI/CD, or observability.
- Familiarity with Mapbox technologies, location data platforms, or mapping-related applications.
- Background in industries that rely on real-time location data such as logistics, transportation, or mobility.
- Previous work in highly regulated environments requiring strict compliance and auditability.
- Contributions to disaster recovery and business continuity planning at scale.
Practical notes
This role is based in the United States and requires eligibility to work in the country without sponsorship. The Cloud Platform team operates across multiple time zones, requiring the ability to work asynchronously and participate in occasional overlapping hours with colleagues in Europe and Japan. Travel may be required periodically for team in-person gatherings, offsites, and stakeholder meetings. Employment is at-will and must comply with Mapbox employment policies and guidelines.