Operations Engineer
Job description
Operations Engineer at Coates Group.
About the role
The Operations Engineer at Coates Group brings a deep specialization in Operations and DevOps practices to the table, ensuring the smooth execution of infrastructure initiatives. This role is defined by a hybrid responsibility model where the hire may be embedded within a specific development squad to provide dedicated application infrastructure assistance or operate more broadly within a centralized Ops Team focusing on deployment and monitoring strategies at an enterprise scale. The position demands a proactive approach to research and analysis, where the engineer evaluates emerging technologies and best practices to align solutions with dynamic program objectives. Success in this role is measured by the ability to translate complex technical concepts into actionable advice for clients and to mentor junior team members through collaborative platform design improvements. The engineer plays a critical part in maintaining the integrity and uptime of a massive global infrastructure, requiring comfort with high-stakes deployments and constant system optimization. This is a strategic position that bridges the gap between development velocity and operational reliability in a fast-paced environment.
Key facts
What you'll do
Orchestrate and manage the deployment lifecycle of over 100 AWS instances distributed across 50 virtual private clouds on a global scale, ensuring architectural integrity and security compliance.
Design, implement, and maintain the management framework for a global fleet of tens of thousands of Linux-powered IoT devices, focusing on scalability and remote reliability.
Monitor complex system infrastructures and analyze alerting data in real-time to drive rapid incident response and guarantee maximum system uptime.
Investigate and dissect existing business workflows to identify bottlenecks and devise innovative methods to integrate these processes into scalable product solutions.
Leverage advanced containerization technologies, specifically Docker, to create robust and portable environments for application deployment and testing.
Develop and deliver comprehensive training materials and conduct formal educational sessions to upskill junior team members on platform operations and infrastructure tools.
Utilize strong scripting capabilities in Bash and Python to automate repetitive tasks, enhance monitoring workflows, and build custom tooling for infrastructure management.
Apply Infrastructure as Code methodologies using tools like CloudFormation to automate the provisioning and management of cloud resources with version control.
Collaborate with cross-functional teams to champion and implement new initiatives that enhance cloud infrastructure resilience, performance, and cost-efficiency.
Perform seamless deployments across diverse runtime environments, including development, user acceptance testing, and production, handling a wide range of device types.
Employ configuration management tools to ensure consistent and reliable server builds, configurations, and updates across the entire technology stack.
Utilize monitoring and observability platforms such as SumoLogic and AWS CloudWatch to visualize trends, troubleshoot issues, and drive data-informed decisions.
Maintain a strong proficiency in a hybrid software stack, balancing serverless functions with microservices to deliver resilient and cost-effective architectures.
Act as a technical leader within a DevOps culture, promoting collaboration between development and operations to streamline the software delivery pipeline.
Requirements
You must possess a highly advanced knowledge of Amazon Web Services, including its core services, security models, and global networking infrastructure.
You must have a solid grasp of advanced containerization concepts, particularly Docker, including orchestration and image optimization techniques.
You must be capable of creating detailed user documentation and training materials for internal platforms and procedures.
You must have a strong background in Linux system administration and network configuration, with Ubuntu being the primary operating environment.
You must demonstrate a strong aptitude for automation and templating, with a firm understanding of Infrastructure as Code principles and tools.
You must have hands-on experience with configuration management tools such as Ansible, Puppet, or Chef.
You must have proven experience with monitoring and observability dashboard tools like SumoLogic and AWS CloudWatch for performance analysis.
You must be able to work effectively in a hybrid software stack environment, managing both microservices and serverless function architectures.
Nice to have
Experience working with IoT device management platforms and firmware update mechanisms.
Knowledge of security and compliance frameworks relevant to cloud infrastructure.
Practical notes
Work hours: Standard business hours.
Travel: Minimal travel required.
Visa: Local candidates preferred.