Platform Engineer
Job description
Platform Engineer - at Bitso.
About the role
You will own the complete incident lifecycle at Bitso, acting as the central owner from detection through resolution and continuous improvement. You will operate confidently in high-pressure production incidents, communicating clearly with senior stakeholders and leadership while keeping systems reliable. You will drive postmortem processes and translate learnings into concrete automation that reduces toil and prevents recurrence. You will champion a culture where asking "how do we make sure this never happens again?" leads to built solutions, not just documentation. If you thrive under pressure, love automation, and want to make a measurable dent in how a high-scale crypto platform operates, this role was designed for you. You will partner closely with cross-functional teams to harden platform reliability and turn incident response into a strategic advantage.
Key facts
What you'll do
- Own and execute on-call shifts end-to-end: acknowledge pages within SLA, declare incidents, assign roles, maintain comms cadence, and drive to resolution.
- Build automation that drives the Sev1/Sev2 postmortem workflow - from scheduling and facilitation reminders to action-item assignment, ownership tracking, and due-date enforcement.
- Leverage AI to identify patterns across incidents and propose systemic fixes: runbook improvements, alert tuning, platform hardening, and process changes.
- Build and extend internal automation to reduce manual toil in detection, escalation, and recovery activities across critical services.
- Navigate Kubernetes clusters with confidence, diagnosing pod-level issues and ensuring workload stability under high transaction volumes.
- Collaborate with product and infrastructure teams to design resilient architectures that minimize single points of failure and improve observability.
- Analyze incident trends to prioritize platform investments that reduce noise, accelerate mean time to resolution, and improve customer trust.
- Partner with engineering and support teams to communicate incident status, timelines, and remediation steps in a clear, structured manner.
- Maintain and evolve runbooks and playbooks so they reflect the latest system behavior and operational best practices.
- Contribute to the development of AI agents or LLM-based workflows that enhance incident analysis and remediation suggestions.
- Measure the effectiveness of reliability initiatives and iterate based on quantitative outcomes and stakeholder feedback.
- Promote best practices across the organization by mentoring peers and elevating incident management standards.
- Work closely with the Incident Management Manager to align on priorities, improve processes, and scale operations responsibly.
- Participate in on-call rotations to maintain hands-on familiarity with production systems and failure modes.
- Champion a mindset where reliability is a first-class feature, not an afterthought, influencing roadmap decisions.
- Engage in continuous learning to keep up with evolving tools, platforms, and incident management techniques.
Requirements
- Proven ability to operate confidently in high-pressure incident scenarios, including communicating clearly with senior stakeholders and leadership while a production issue is live.
- Hands-on experience with Kubernetes - comfortable deploying, debugging, and navigating pod-level issues.
- Solid understanding of CI/CD pipelines and modern DevOps practices.
- Software development background in any language; ability to read, write, and debug code is essential (Python or Java experience is a plus).
- Strong automation mindset: you identify repetitive toil and your first instinct is to eliminate it, not absorb it.
- Experience building or working with AI agents or LLM-based workflows is highly desirable.
- Strong interpersonal and written communication skills.
- Self-directed learner who doesn't need a fully defined path to start contributing.
- Fintech or crypto industry background is a plus - familiarity with the domain vocabulary accelerates onboarding and incident triage.
- Comfortable working in a fast-paced, dynamic environment where priorities can shift quickly.
- Willingness to follow and contribute to strict incident management processes and communication protocols.
- Demonstrated ownership of complex systems and an inclination to drive them toward robustness.
- Commitment to professional growth and collaboration within a distributed, diverse team.
- Ability to manage multiple responsibilities and maintain focus during critical incident windows.
Nice to have
- Experience building or working with AI agents or LLM-based workflows.
- Background in fintech or cryptocurrency domains.
Practical notes
- This role is based in Latin America and follows the standard working hours and practices defined by Bitso.
- Travel is not required for this position.
- Visa support is available for eligible candidates as needed.
- Candidates must be able to start within the timeframe specified in the source information.