DevOps Platform Engineer
EarnInMexico1w ago
Job description
About the role
EarnIn is looking for a DevOps Platform Engineer to help build and maintain the internal platform that supports their product delivery. This role is located in Mexico City, Mexico and offers remote work arrangements within Mexico. The engineer will be responsible for designing reliable infrastructure, automating deployment processes, and ensuring the stability of production systems. You will partner with engineering teams to create scalable solutions that enable fast and safe software releases.
Key facts
What you'll do
- Design and implement CI/CD pipelines to automate build, test, and deployment workflows across multiple microservices and applications.
- Manage containerized workloads using orchestration platforms to ensure high availability, fault tolerance, and efficient resource utilization in production.
- Monitor system performance metrics and set up comprehensive alerting mechanisms to detect anomalies and resolve incidents before they escalate.
- Collaborate closely with software engineers to define infrastructure requirements, capacity planning, and deployment strategies for new product features.
- Maintain and evolve infrastructure-as-code repositories to keep all environment configurations versioned, auditable, and fully reproducible across stages.
- Troubleshoot production incidents by analyzing distributed logs, application metrics, and request traces to quickly identify and resolve root causes.
- Optimize cloud resource allocation and usage patterns to balance cost efficiency with performance targets and reliability requirements.
- Implement security best practices throughout the platform, including identity and access management, network segmentation, and compliance validation checks.
- Develop internal automation tooling and reusable scripts to reduce manual operational overhead and improve team productivity over time.
- Participate in structured on-call rotations to provide timely first-response support for production systems and service-level agreement commitments.
- Build and maintain self-service platforms that allow development teams to deploy and manage their own applications independently.
- Conduct post-incident reviews and document findings to drive continuous improvement in system reliability and operational procedures.
Requirements
- Proven experience with Linux system administration and networking fundamentals in production environments, including service management and configuration.
- Solid familiarity with container technologies and orchestration platforms for deploying, scaling, and managing distributed application workloads reliably.
- Strong knowledge of CI/CD concepts with hands-on experience in configuring, maintaining, and improving automated deployment pipelines.
- Practical understanding of cloud computing platforms and their core services for infrastructure provisioning, networking, and storage management.
- Demonstrated ability to write scripts in at least one scripting or programming language to automate operational tasks effectively.
- Excellent problem-solving skills with the capacity to diagnose complex system failures and communicate findings clearly to stakeholders.
- Experience working in a collaborative, cross-functional team environment where clear communication and shared ownership are valued practices.
- Familiarity with version control systems and code review workflows as part of a team-based software development culture.
- Comfort with ambiguity and the ability to prioritize competing tasks in a fast-paced, rapidly evolving technology environment.
- Track record of documenting operational procedures and knowledge sharing with teammates to reduce single points of failure.
Nice to have
- Exposure to observability platforms and distributed tracing tools for monitoring microservices architectures and understanding service dependencies.
- Experience with policy-as-code frameworks for enforcing governance, compliance, and security standards across cloud infrastructure environments.
- Background in financial technology or payments infrastructure with awareness of regulatory requirements and data protection considerations.
- Contributions to open-source projects or community-driven infrastructure tooling initiatives that have benefited broader engineering organizations.
- Familiarity with chaos engineering practices and resilience testing methodologies for validating system behavior under failure conditions.
Skills & tools
- Linux operating systems and command-line administration for server management, troubleshooting, and performance tuning in production.
- Container runtimes and orchestration platforms for deploying, scaling, and managing containerized application workloads efficiently.
- CI/CD tooling and pipeline frameworks for building automated build, test, and deployment workflows across environments.
- Cloud infrastructure providers and their core services for provisioning compute, storage, and networking resources on demand.
- Scripting and general-purpose programming languages for writing automation, internal tooling, and operational utilities.
- Monitoring, logging, and observability solutions for gaining visibility into system health, performance, and incident response.
Practical notes
- This role operates in a remote capacity within Mexico, with the primary office location in Mexico City.
- The position is full-time with standard business hours, subject to occasional on-call duties for production incident response.
- Candidates should expect to participate in regular team ceremonies, sprint planning, and cross-functional collaboration meetings consistently.
- The hiring process includes multiple technical interviews focused on infrastructure design, automation expertise, and real-world problem-solving scenarios.
- This is a permanent, full-time position with opportunities for professional growth and skill development over time.