Senior Platform Engineer
Job description
About the role
The NitroAI team is seeking a dedicated engineer to manage the analytics infrastructure used by our business consulting division. In this capacity, you will serve as the primary platform owner and operate an AWS and Kubernetes environment with full autonomy to drive technical strategy and execution. You will be responsible for ensuring that the underlying systems remain robust, efficient, and aligned with evolving business needs. This role requires a proactive mindset focused on automation, reliability, and continuous improvement. You will partner closely with consulting teams to translate analytical requirements into scalable platform capabilities. The position offers significant influence over the technical direction and day-to-day operations of critical infrastructure. Your work will directly impact the performance and stability of analytics workloads supporting our enterprise customers.
Key facts
What you'll do
- Manage the reliability, performance, and cost efficiency of the AWS and Kubernetes platform through vigilant oversight and optimization.
- Direct GitOps deployment workflows, including configuration management and policy synchronization to ensure consistency across environments.
- Oversee the end-to-end CI/CD lifecycle, including GitLab pipeline maintenance, Docker runner management, and environment promotion procedures.
- Manage incident response activities, including troubleshooting complex issues, executing recovery plans, and documenting post-mortem analyses.
- Build and maintain observability, monitoring, and alerting systems specifically tailored for analytics consulting workloads and use cases.
- Enforce security standards, access controls, and secrets management to maintain a strong security posture for all platform components.
- Collaborate with cross-functional stakeholders to design and implement infrastructure solutions that meet scalability and availability targets.
- Automate routine operational tasks to reduce manual effort and improve efficiency across the platform lifecycle.
- Participate in on-call rotations to provide timely support and guidance for production issues as they arise.
- Mentor junior engineers on best practices for cloud infrastructure, containerization, and deployment strategies.
- Evaluate emerging technologies and tools to identify opportunities for enhancing the platform capabilities and developer experience.
- Ensure all infrastructure changes comply with organizational policies and regulatory requirements relevant to data handling.
- Contribute to architectural decision-making processes and long-term roadmap planning for the analytics platform.
- Drive continuous improvement initiatives aimed at reducing costs, increasing resilience, and enhancing system performance.
Requirements
- 5+ years of platform engineering experience building distributed, scalable systems that handle real-world production workloads.
- 3+ years of production experience with AWS and Kubernetes, specifically EKS, RBAC, and IRSA, demonstrating a deep understanding of these services.
- Expert knowledge of GitLab CI/CD, including shell and Docker runner management and Bash scripting, to automate and manage pipelines effectively.
- Production experience with ArgoCD and GitOps workflows across multiple environments, ensuring reliable and repeatable deployments.
- Proficiency in Terraform for Infrastructure as Code, enabling consistent and version-controlled infrastructure provisioning.
- AWS expertise covering Lambda, ALB, WAF, CloudFront, Cognito, Secrets Manager, Route53, KMS, ECR, and VPC networking to design and manage cloud solutions.
- Understanding of security baselines and compliance-focused infrastructure to meet enterprise standards and regulatory obligations.
- Strong problem-solving skills and the ability to troubleshoot complex issues in distributed systems environments.
- Excellent communication skills to collaborate effectively with technical and non-technical stakeholders.
- Commitment to following established processes while also identifying opportunities for process improvement and innovation.
- Willingness to work within the framework of a public benefit corporation and align with its mission and values.
- Ability to manage multiple priorities in a fast-paced environment while maintaining attention to detail.
- Readiness to engage in professional development and stay current with industry trends and best practices.
- Capability to work independently with minimal supervision while contributing effectively to team objectives.
Nice to have
- Experience with Apache Airflow via MWAA to manage complex workflow orchestration and scheduling requirements.
- Proficiency with Claude Code to leverage advanced AI-assisted development and coding practices.
- Familiarity with Okta or SSO integration to streamline authentication and access management processes.
- Experience using Serverless Framework for Lambda pipelines to simplify the deployment and management of serverless applications.
- Knowledge of Databricks Asset Bundles for job and notebook deployment to enhance data engineering workflows.
- Experience with Grafana for observability to create detailed dashboards and monitor system health effectively.
- Background in SOC-2 compliance for data engineering to ensure data security and governance standards are met.
Practical notes
- Veeva is a public benefit corporation, reflecting a commitment to social responsibility and stakeholder value.
- The hiring process includes a personality assessment within 3 days of application, followed by a manager interview, a case study, and a final leadership interview.
- Benefits include medical, dental, vision, life insurance, retirement programs, and flexible PTO to support work-life balance.
- This role is based in Massachusetts
- Boston, but follows a Work Anywhere policy, allowing for flexibility in work location.
- Contact talent_accommodations@veeva.com for disability-related assistance to ensure inclusive hiring practices.
- No specific hours are stated, implying standard full-time expectations within the role.
- Travel requirements are not specified, suggesting the position is primarily focused on local or remote work as per company policy.
- Visa sponsorship information is not provided, indicating standard eligibility criteria will apply.
- Deadlines for application review are not outlined, so candidates should apply promptly to ensure full consideration.