DevOps II
Job description
About the role
This position is for a DevOps Engineer at Storable focused on architecting and owning the end-to-end reliability of containerized workloads on Kubernetes clusters in production. The role involves designing and implementing CI/CD pipelines for application deployment with an emphasis on stability and traceability. You will automate infrastructure provisioning and configuration management using code-driven approaches across the technology stack. Responsibilities include monitoring system performance, reliability, and availability metrics to ensure service continuity. You will collaborate closely with development and QA teams to enable seamless and controlled releases. The position also entails troubleshooting and resolving complex infrastructure and deployment issues while maintaining detailed runbooks. Additionally, you will support AI/ML model deployment and basic MLOps workflows to bridge data science and production operations.
What you'll do
You will design, implement, and manage CI/CD pipelines for application deployment with version control and rollback capabilities. You will work with containerization and orchestration tools, especially Kubernetes, to define deployments, services, and ingress resources. Automating infrastructure provisioning and configuration management using declarative code and templates for repeatability will be a core responsibility. You will monitor system performance, reliability, and availability through metrics, logs, and alerting mechanisms. Collaboration with development and QA teams for seamless releases by integrating testing and validation stages will be essential. You will troubleshoot and resolve infrastructure and deployment issues by analyzing logs, traces, and system behavior. Supporting AI/ML model deployment and basic MLOps workflows, including model packaging and serving infrastructure, will be required. Implementing security best practices in pipeline design, such as least privilege access and secret management, will be part of the role. Optimizing resource utilization on clusters by adjusting requests, limits, and autoscaling configurations will be expected. Maintaining documentation for environments, deployment procedures, and operational runbooks is required. Participation in on-call rotations to respond to incidents and support rapid resolution will be necessary. Evaluating and recommending new tools and technologies to improve the efficiency of the DevOps lifecycle will be important. Ensuring compliance and governance policies are embedded into infrastructure as code and pipeline definitions will be required. Driving continuous improvement by measuring deployment frequency, failure rate, and lead time for changes will be part of the responsibilities.
Requirements
The role requires 4+ years of experience in DevOps or related roles with a proven track record of delivering reliable systems. Strong hands-on experience with Kubernetes, including cluster operations, networking, and storage concepts, is essential. Proficiency in scripting languages such as Python, Bash, or Shell to automate tasks and build tooling is required. Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or similar platforms is necessary. Knowledge of cloud platforms like AWS, Azure, or GCP, including core services and networking, is required. Experience with containerization tools like Docker, including image builds, registry management, and security, is needed. Understanding of Infrastructure as Code using tools such as Terraform, CloudFormation, or equivalent is required. The ability to work in a fast-paced environment and manage multiple priorities with clear communication is essential.
Nice to have
Basic knowledge of AI/ML concepts and model deployment practices to support data science initiatives is preferred. Familiarity with MLOps tools and workflows for managing model lifecycles and experiments is beneficial. Experience with monitoring tools such as Prometheus, Grafana, or the ELK stack for visibility is advantageous. Knowledge of security best practices in DevOps, including vulnerability scanning and policy enforcement, is preferred.
Practical notes
This engagement location is Hyderabad, Telangana, India, and may require local travel within the defined work area. Visa requirements for international candidates will be handled per local regulations and corporate policy. Please apply with updated documentation and relevant project examples that match the outlined responsibilities. The deadline for submission of applications is as specified in the official portal or as communicated by the hiring team.