Platform Infrastructure Engineer
Job description
About the role
At Menlo Security, our mission is enabling the world to connect, communicate and collaborate securely without compromise. This mission has become even more critical in the context of COVID-19, and we are actively scaling to support customers that include Fortune 500 companies, 9/10 of the largest global banks, and the Department of Defense. As we grow from 400 employees to the next phase of our journey, we are seeking talent that is ethical, hyper-organized, fanatical about seeing things through to completion, service-oriented, and humble enough to both give and receive feedback. This role is backed by investors such as Vista Equity Partners, General Catalyst, JPMC, American Express, HSBC, and Ericsson Ventures, and is integral to our continued growth.
The Platform Infrastructure Engineering team builds and operates the infrastructure that powers our global platform. In this role, you will own the design, implementation, and day-to-day operation of the core infrastructure platform that enables our customers to connect to the Internet without compromise. You will extend our cloud-native environment on Google Kubernetes Engine while maintaining robust multi-region resilience, eliminate operational toil, and partner closely with security, product, compliance, and engineering teams to translate business requirements into reliable infrastructure solutions. This role is critical to ensuring that Menlo's platform remains secure, observable, and always available.
What you'll do
- Implement, deploy, and maintain virtual machine and Kubernetes infrastructure on Google Cloud Platform and AWS across dozens of clusters spanning development, staging, and production environments in multiple regions.
- Build and maintain Infrastructure as Code using Terraform modules and Spacelift, provisioning networking, compute, storage, and security components while implementing multi-layer configuration management workflows.
- Implement and maintain observability solutions using Grafana Cloud, Prometheus/Mimir, and OpenTelemetry collectors, designing dashboards and alerting rules that provide comprehensive, end-to-end visibility into all platform components.
- Manage certificate lifecycle, DNS automation, ingress controllers, and service mesh networking with Cilium to ensure secure and reliable traffic flow across the platform.
- Partner with peers and across Engineering, Product, Compliance, and Security teams to align on requirements and consult on capacity planning, disaster recovery, and architectural decisions.
- Identify and eliminate operational toil by writing scripts, building CI/CD pipelines, and leveraging AI-assisted development and code-review tools, including Gemini Code Assist, to build and troubleshoot infrastructure code efficiently and safely.
- Participate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviews to improve reliability and reduce recurrence.
- Utilize LLM-based tooling to accelerate infrastructure code development, debugging, and troubleshooting as part of the standard engineering workflow.
Requirements
- Hold a Bachelor's degree in Computer Science, a related technical field, or demonstrate equivalent practical experience.
- Demonstrate proficiency in common programming and scripting languages, with a strong emphasis on Python, Bash, and Go to automate complex infrastructure tasks and solve operational challenges.
- Show a clear understanding of network topologies, communication protocols (including TCP/IP, HTTP/S, UDP, and TLS), and enterprise-grade connectivity solutions.
- Exhibit deep Kubernetes expertise, including cluster administration, Role-Based Access Control (RBAC), networking, workload management, and troubleshooting in production-grade environments.
- Prove hands-on experience with Terraform for infrastructure provisioning and management, ensuring that cloud resources are defined and deployed safely, consistently, and securely.
- Maintain a thorough understanding of Google Cloud Platform services, including GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, IAM, Gemini Code Assist, and Workload Identity.
- Demonstrate a clear ability to use LLM-based code-assist and code-review tools to effectively build, troubleshoot, and maintain infrastructure code as part of standard engineering practices.
- Operate with the core values outlined in our mission: be ethical, hyper-organized, fanatical about execution, service-oriented, humble enough to absorb feedback, and confident enough to provide coaching to peers.
Preferred / Nice to Have
- Experience with GitOps methodologies and tools to further streamline deployment workflows and improve operations reliability.
Compensation and Benefits
Menlo Security is committed to competitive compensation and comprehensive benefits. While specific details are provided during the hiring process, the total rewards package is designed to reflect our investment in top-tier talent and our growth stage. As a fully remote-friendly role based in AMER
Canada, we ensure that all eligible team members have the tools and support needed to perform their best.