Senior Infrastructure Engineer
Job description
About the role
Join a dedicated and focused team of two engineers responsible for maintaining and advancing the platform that empowers millions of creators worldwide. This role is ideal for someone who enjoys building developer tools as products and ensuring the reliability and stability of core systems. You will play a key part in modernizing our infrastructure, improving developer workflows, and supporting the platform's scalability and performance. The position offers an opportunity to work on a variety of cloud technologies and infrastructure practices, contributing to a highly impactful product used by a global community of creators. As part of the team, you will collaborate closely with engineering, product, and design teams to ensure the platform remains reliable, efficient, and scalable, while also focusing on automating routine tasks and enhancing developer experience.
Key facts
What you'll do
- Ensure the consistent operation, stability, and performance of our production platform, including managing Amazon EKS clusters and CI/CD pipelines to support rapid and reliable deployments.
- Develop and refine advanced deployment strategies such as progressive delivery to make software releases safer, faster, and more predictable, reducing downtime and rollback risks.
- Build and maintain developer tools that function as well-supported products, significantly improving the local development environment and streamlining workflows for engineers.
- Automate routine operational tasks using AI tools to reduce manual toil, increase efficiency, and allow engineers to focus on more complex and impactful challenges.
- Manage and update our technology stack, including runtimes, Kubernetes configurations, Helm charts, and Terraform scripts, ensuring they are maintainable, scalable, and secure.
- Improve the cost-effectiveness of our cloud infrastructure by optimizing resource usage, increasing visibility into costs, and implementing best practices for cloud spend management.
- Collaborate closely with engineering, product, and design teams to identify infrastructure needs, improve platform reliability, and implement new features or improvements that enhance the overall developer and user experience.
- Support incident response efforts, conduct post-mortem analyses, and implement improvements to alerting and monitoring systems to prevent future issues.
- Contribute to building internal developer tools such as command-line interfaces (CLIs) and local development environments, making it easier for engineers to develop, test, and deploy code efficiently.
- Advocate for best practices in security, reliability, and operational excellence, ensuring infrastructure adheres to industry standards and internal policies.
- Participate in regular infrastructure reviews, capacity planning, and documentation efforts to maintain a high standard of operational readiness.
- Use modern AI tools to assist in debugging, documentation, and reducing manual work, leveraging automation to improve system reliability and developer productivity.
- Make pragmatic build-versus-buy decisions based on operational costs, maintainability, and team capacity, ensuring the infrastructure remains cost-effective and scalable.
- Share knowledge and contribute to a culture of continuous improvement by writing documentation, blog posts, or sharing insights with the team and broader community.
Requirements
- Proven experience supporting live systems in infrastructure, SRE, DevOps, or platform engineering roles, with a focus on reliability and operational excellence.
- Hands-on experience managing production Kubernetes environments, including creating Helm charts, configuring autoscaling, resource management, and troubleshooting issues.
- Strong knowledge of AWS cloud services, including EC2, S3, SQS, SNS, ECR, IAM, and load balancers, with practical experience in designing and maintaining cloud infrastructure.
- Proficiency with Terraform, including modular design, maintainable code, and managing infrastructure as code in a production environment.
- Experience operating CI/CD pipelines and GitOps workflows, with a focus on reliable deployment, rollback strategies, and automation.
- Demonstrated ability to manage incidents, conduct thorough post-mortems, and implement improvements to alerting and monitoring systems.
- Experience in building internal developer tools such as CLIs, local development environments, or automation scripts that improve developer productivity.
- Ability to work effectively in a remote, asynchronous environment, communicating clearly, providing context, and collaborating across time zones.
- A strong drive to identify recurring issues, implement iterative improvements, and reduce operational toil.
- Familiarity with modern AI tools used for debugging, documentation, and automation to enhance system reliability and developer workflows.
- Ability to evaluate operational costs pragmatically, making informed build-versus-buy decisions that align with team goals and budget constraints.
- Personal interest in creating and sharing work, whether through writing, coding, or other forms of content, to contribute to the broader community and team knowledge base.
Nice to have
- Familiarity with the social media management industry or experience as a Buffer user, understanding the platform's specific needs and workflows.
- Experience with observability and monitoring tools such as Datadog or Sentry, with a focus on cost-aware logging, metrics, and alerting strategies.
- Prior experience working with Cloudflare products, including Workers, Zero Trust, and DNS management, to enhance security and performance.
- Contributions to open-source projects, particularly Terraform modules or tools that support infrastructure automation and management.
- Knowledge of additional cloud providers or services beyond AWS, such as Google Cloud Platform (GCP), to support multi-cloud strategies.
- Experience with security best practices in cloud infrastructure, including network segmentation, access controls, and compliance considerations.
Skills & tools
- AWS, GCP, Cloudflare
- EC2, EKS, S3, SQS, SNS, ECR, IAM, ALB, BigQuery
- Terraform
- Kubernetes, Helm, KEDA, ArgoCD, Argo Rollouts
- GitHub Actions
- Datadog, Sentry, Incident.io, PagerDuty
- MongoDB, Elasticsearch, Redis
- Node.js, TypeScript, Python, PHP
- CloudFront, VPC peering, OctoDNS
- OrbStack
Practical notes
- This is a fully remote position, allowing you to work from anywhere.
- Buffer team members are encouraged to travel once or twice a year for team events, which are optional but appreciated for team bonding.
- Applications are reviewed carefully by a small, dedicated team committed to finding the right fit.
- The hiring process includes multiple stages: application review, interviews with the hiring manager, a take-home exercise, technical and leadership interviews, a final interview, and a paid collaboration period on a real project to assess practical skills.
- Buffer values diversity and is committed to creating an inclusive environment; applications from individuals with varied backgrounds are encouraged.
- By applying, you consent to Buffer processing your personal data for recruitment purposes, in accordance with applicable privacy policies.