
Senior DevOps Engineer
Job description
Senior DevOps Engineer at Aiops Group.
About the role
Aiops Group is seeking a Senior DevOps Engineer to join their growing team in Sofia, Bulgaria. This role is a full-time employee position focused on supporting and advancing the company's infrastructure and deployment practices across all engineering teams. The engineer will work closely with development and operations stakeholders to ensure reliable, scalable, and secure systems in production. The position involves shaping the technical direction of platform operations and contributing to long-term architectural decisions that impact the entire organization. The Senior DevOps Engineer will also mentor junior team members and contribute to the growth of the engineering culture at Aiops Group.
Key facts
What you'll do
Own the design and maintenance of CI/CD pipelines that support application delivery across multiple environments and stages. Monitor system performance and availability metrics, responding to incidents promptly and implementing preventive measures to reduce future occurrences. Collaborate with development teams to define infrastructure requirements for new features, services, and product releases. Manage and automate provisioning of cloud resources to support growing operational demands and expanding service portfolios. Conduct root cause analysis for production issues and document findings thoroughly for the broader engineering team. Implement infrastructure-as-code practices to ensure consistent and repeatable environment configurations across development, staging, and production. Evaluate and recommend tools and platforms that improve deployment speed, system reliability, and overall operational efficiency. Participate in on-call rotations and provide tier-three support for critical production systems during high-severity incidents. Drive standardization of operational procedures, runbooks, and documentation practices across all engineering teams and projects. Partner with security stakeholders to ensure compliance and best practices are followed in infrastructure management and configuration. Review and optimize existing deployment workflows to reduce lead times and increase the frequency of safe releases. Troubleshoot and resolve infrastructure-related issues in coordination with development and QA teams to minimize service disruptions.
Requirements
Demonstrated experience in a DevOps or site reliability engineering role within a professional technology environment. Strong understanding of continuous integration and continuous delivery methodologies, including pipeline design, implementation, and maintenance. Familiarity with scripting and automation languages commonly used for infrastructure management, configuration, and orchestration tasks. Experience managing containerized workloads and orchestration platforms in production environments at scale and under load. Knowledge of monitoring, logging, and observability practices for distributed and microservice-based systems in production. Ability to work effectively in a collaborative team environment with cross-functional engineering partners and stakeholders. Proven track record of improving system reliability, reducing operational incidents, and driving meaningful automation efforts. Comfortable working in a fast-paced setting with shifting priorities, competing deadlines, and changing business requirements. Strong communication skills for presenting technical concepts to both engineering and non-technical audiences clearly. Experience with version control and code review processes to maintain quality and consistency in infrastructure code.
Nice to have
Experience with multi-cloud or hybrid infrastructure deployments across different cloud providers and geographic regions. Background in working with agile or Scrum development methodologies, including sprint planning and retrospective ceremonies. Familiarity with policy-as-code frameworks and governance automation for cloud resource management and compliance enforcement. Exposure to service mesh technologies and advanced networking concepts in container orchestration and microservice environments.
Skills & tools
Proficiency in scripting languages such as Python, Bash, or Go for automation and infrastructure management tasks. Experience with configuration management tools like Ansible, Puppet, or Chef for automating infrastructure provisioning and updates. Working knowledge of containerization technologies including Docker and container orchestration platforms in production settings. Familiarity with version control systems, particularly Git, for managing infrastructure and pipeline code repositories effectively. Understanding of cloud provider services, resource management, and cost optimization strategies in production contexts. Comfort with monitoring and alerting stacks including tools like Prometheus, Grafana, or similar observability platforms.
Practical notes
This is a full-time, on-site position based in the Sofia, Bulgaria office location. The hiring process may include technical interviews, system design discussions, and cultural fit assessments with the engineering team. Candidates should be prepared to discuss past infrastructure projects and their specific contributions to those initiatives. Aiops Group values continuous learning and encourages engineers to attend conferences, participate in internal knowledge-sharing sessions, and share knowledge with their peers regularly.