Senior Support Engineer
Job description
About the role
This Senior Support Engineer role at Alternative Payments represents a pivotal opportunity to own the end-to-end lifecycle of critical production issues in a high-growth FinTech environment. You will serve as the primary technical escalation leader, bridging the gap between our engineering teams and both internal stakeholders and external clients to ensure seamless support delivery. The position demands hands-on troubleshooting across distributed systems, infrastructure, and application layers to resolve complex incidents efficiently. You will design and drive the implementation of comprehensive monitoring, automation, and operational playbooks that define our support standards. Success in this role means taking full ownership of problem resolution from initial detection through to permanent solution delivery. You will leverage cutting-edge AI-powered tools to streamline issue diagnosis and automate repetitive workflows. Ultimately, you will be instrumental in building strategic support initiatives that directly enhance client trust and reinforce our commitment to operational excellence.
Key facts
What you'll do
- Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions, ensuring minimal production downtime.
- Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
- Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes, reducing manual overhead and improving response times.
- Collaborate with cross-functional teams to deliver solutions, optimize support processes, and implement scalable monitoring systems that enhance overall reliability.
- Take ownership of complex production issues and provide technical expertise across distributed systems, infrastructure, and application layers to maintain stability.
- Develop and maintain automation scripts and tools to reduce manual intervention and improve system reliability, ensuring processes are efficient and repeatable.
- Propose improvements and help establish best practices for incident response, monitoring, log management, and proactive issue prevention to mitigate future risks.
- Bridge communication between engineering teams, customer experience, and clients during critical incidents and implementation challenges to ensure alignment and transparency.
- Analyze recurring problems and develop strategic solutions to prevent future occurrences, driving long-term improvements in system robustness.
- Champion a culture of continuous improvement by sharing knowledge, mentoring peers, and contributing to the evolution of support methodologies.
- Partner with product and engineering teams to provide feedback from the support frontline, influencing product decisions and roadmap priorities.
- Conduct regular reviews of support metrics and performance indicators to identify trends and drive operational efficiency.
- Implement and refine observability strategies to gain deeper insights into system behavior and enhance troubleshooting capabilities.
- Lead post-incident review processes, ensuring that lessons learned are documented and actionable improvements are implemented promptly.
- Act as a subject matter expert within the organization, providing guidance on best practices for support engineering and technical operations.
Requirements
- 3-5 years of experience in technical support, SRE, Forward Deployed Engineer, or full-stack engineering roles within technology or FinTech environments.
- Strong expertise with observability and monitoring tools such as DataDog, ELK Stack (Elasticsearch, Logstash, Kibana), Prometheus, Grafana or similar platforms for real-time insights.
- Solid infrastructure and DevOps knowledge including infrastructure as a code (IaC), containerization (Docker/Kubernetes) and cloud-native architecture principles.
- Strong skills in AWS infrastructure, observability tools, database management, and CI/CD pipelines to support scalable and reliable systems.
- Solid experience in troubleshooting distributed systems and production environments under pressure and strict uptime requirements.
- Strong communication skills (especially in English), to collaborate effectively across engineering teams, customer success, and external clients with clarity and professionalism.
- A proactive mindset with the ability to solve complex technical problems, work with SLA/SLO frameworks, and drive automation initiatives that improve efficiency.
- Experience with incident management processes and high-intensity support environments where quick decision-making is critical.
- Eligibility to work remotely from Brazil is a mandatory requirement for this position.
- A demonstrated ability to take ownership of issues end-to-end and drive resolutions without constant supervision.
- Commitment to maintaining the highest standards of service delivery and customer satisfaction in all interactions.
- Willingness to participate in on-call rotations and support critical incidents as they arise, ensuring business continuity.
- Adherence to established security and compliance protocols when handling client data and system information.
- Alignment with company values of collaboration, ownership, and a customer-obsessed mindset.
Nice to have
- Experience in FinTech, payments, startup, or scale-up environments where rapid iteration and adaptability are essential.
- Familiarity with multiple programming languages, frontend/backend technologies, and modern observability stacks to understand technical nuances.
- A track record of implementing automation solutions, improving support processes, or driving operational improvements that have measurable impact.
- Comfort working in fast-paced, dynamic, and high-impact environments with strict SLAs and aggressive growth targets.
- Previous experience contributing to open source projects or technical communities that demonstrate thought leadership.
- Exposure to AI-powered monitoring and diagnostic tools that enhance support workflows.
- Knowledge of regulatory and compliance considerations in the payments industry.
Practical notes
This position is fully remote and available exclusively to candidates eligible to work in Brazil. The role operates within standard business hours but may require flexibility for critical incident response. Travel is not required for this position. Visa sponsorship is not applicable due to remote eligibility in Brazil. There are no external deadlines for application, but early submission is encouraged to secure consideration in a growing hiring pipeline.