Senior DevOps Engineer, AI & Applications
Job description
.
About the role
Every AI feature we ship touches thousands of GPUs. The Senior DevOps Engineer will build the release engineering backbone - CI/CD pipelines, automated testing gates, one-click deployments with instant rollback - that lets Firmus scale fast and responsibly. You are the bridge between engineering and operations: setting Firmus standards for how code gets to production, mentoring the team on deployment safety, and driving a blameless culture when things go wrong. Ship safely. Ship often. Ship at scale. You will own the end-to-end reliability of our deployment lifecycle, ensuring that changes flow predictably from commit to cluster. You will design and guard the workflows that let data scientists and engineers move quickly without risking production stability. Your work will directly influence how Firmus responds to customer needs and regulatory expectations. You will continuously refine the pipeline to reduce risk and increase confidence with every release.
Key facts
What you'll do
Design and maintain team-wide CI/CD pipelines (Jenkins, GitHub Actions, ArgoCD, or equivalent) with automated testing gates, artifact management, and deployments aligned with GPU cluster standards.
Implement release engineering best practices: repeatable releases, GitOps workflows, automated rollback, and change management procedure.
Build and manage test infrastructure: environment provisioning, data seeding, long-running job validation (especially for distributed training templates and multi-node job submissions).
Establish engineering protocols and standards: repo organization, PR templates, code quality gates, dependency scanning, static analysis.
Partner with infra teams to ensure AI product features deployment practices meet compliance and security standards for massive GPU clusters.
Mentor team on testing strategies, deployment safety, and incident response procedures.
Integrate testing frameworks into CI pipelines, including unit, integration, and end-to-end tests tailored to AI workloads.
Define and evolve artifact management, versioning strategies, and rollback procedures to support safe and traceable deployments.
Optimize CI/CD performance and cost efficiency as GPU workloads and team size scale.
Champion observability and logging for deployment pipelines to detect and diagnose issues before they affect production.
Drive improvements in deployment frequency, lead time for changes, and recovery time while maintaining high reliability.
Collaborate with security and compliance to ensure deployment practices adhere to organizational policies and external regulations.
Support on-call responsibilities for release-related incidents and contribute to runbooks and post-incident reviews.
Promote a culture of continuous learning by documenting patterns, automating guardrails, and sharing knowledge across engineering teams.
Enable infrastructure-as-code patterns for test and production environments to ensure consistency and reproducibility.
Requirements
5-7 years of CI/CD engineering, release engineering, or DevOps experience.
Deep expertise in GitHub Actions, GitLab CI, ArgoCD, or Jenkins with multi-stage pipeline design and testing gate implementation.
Strong automation scripting (Python, Go, or Bash) for build orchestration and environment templating.
Strong Kubernetes fundamentals (hands-on): deep understanding of Pod lifecycle and failure modes (Pending/Running/CrashLoopBackOff/Evicted), Deployments/ReplicaSets, Jobs/CronJobs, Services/Ingress, and how these primitives behave under load and during rollouts.
Config & secret management: practical experience designing and operating ConfigMaps and Secrets (including secret rotation patterns), with strong hygiene around least privilege, auditability, and preventing credential leakage into logs/artifacts.
Safe rollout patterns: proven experience implementing and operating safe rollout strategies (rolling updates, canary, blue/green), readiness/liveness/startup probes, PodDisruptionBudgets, and rollback procedures - ensuring zero/low-downtime deployments for customer-facing services.
Deployment safety & debugging: ability to debug common Kubernetes rollout issues end-to-end (bad probes, misconfigured resources/limits, image pull failures, secret/config drift, node pressure/evictions) and convert learnings into automated CI/CD gates and runbooks.
Familiarity with artifact management, versioning strategies, and rollback procedures.
Experience integrating testing frameworks into CI pipelines (unit, integration, end-to-end).
8+ years of professional experience in software engineering, DevOps, or related roles.
Proven track record of maintaining and scaling CI/CD systems in production environments.
Strong understanding of containerization, networking, and storage concepts in the context of AI workloads.
Experience working in regulated or compliance-conscious environments is a strong asset.
Excellent written and verbal communication skills for cross-team collaboration and mentoring.