Infrastructure Stability Architect
Job description
About the role
Join our team to transform how we manage system reliability by shifting from reactive fixes to proactive, end-to-end risk identification. You will partner with cross-functional squads to uncover latent hazards in distributed systems before they escalate into customer-impacting events. This role owns the definition of architectural guardrails that preserve continuity for our global crypto trading and wallet products. You will evangelize platform thinking by converting fragile components into resilient, observable, and maintainable services. The position requires a balance between deep technical craftsmanship and strategic influence over long-term reliability outcomes. You will mentor engineers on stability-minded design while ensuring standards are consistently implemented across services. Your work will directly affect the integrity of high-availability infrastructure that underpins mission-critical financial flows. Through continuous risk assessment, you will drive measurable improvements in system robustness and operational maturity.
Key facts
What you'll do
- Systematically probe production behavior and non-production environments to surface performance constraints and architectural weaknesses.
- Map complex service interactions to clarify dependencies and identify single points of failure that threaten continuity.
- Define and evolve architectural patterns that align with reliability objectives, security considerations, and cost efficiency.
- Evaluate emerging infrastructure tools and frameworks, then recommend those that strengthen our stability engineering practices.
- Collaborate with product and platform teams to embed stability criteria into feature and service design lifecycles.
- Construct reference implementations and blueprints that guide engineering teams toward standardized, resilient system designs.
- Quantify the impact of reliability initiatives through metrics, logs, and traces to validate improvements and prioritize efforts.
- Streamline incident response processes by improving runbooks, detection logic, and failure-handling strategies.
- Partner with infrastructure vendors and internal stakeholders to resolve cross-organizational reliability challenges.
- Champion automation that reduces manual toil and increases consistency in environment management and deployment workflows.
- Conduct architecture reviews to ensure that new designs meet scalability, durability, and operability standards.
- Drive the adoption of SRE practices, balancing service-level objectives with realistic delivery constraints.
- Translate business risk into technical requirements so that engineering teams can implement effective protective measures.
- Act as a technical authority on stability matters, providing clear rationale for architectural decisions and trade-offs.
Requirements
- Hold a Bachelor degree or higher in Computer Science, Engineering, or a closely related technical field.
- Bring a minimum of 7 years of professional experience in software development, systems architecture, or SRE roles.
- Demonstrate proficiency in Java, Go, or C++ along with strong knowledge of DevOps and SRE principles in practice.
- Show hands-on experience with middleware technologies such as Nginx, Redis, ElasticSearch, Kafka, MySQL, Docker, and Kubernetes.
- Prove ability to design business architectures and build complete systems from the ground up, considering data flow, failure modes, and scaling constraints.
- Exhibit sound judgment in selecting technologies and patterns that balance innovation with operational risk.
- Display strong analytical skills to deconstruct complex problems and model system behavior under load and failure conditions.
- Maintain a disciplined approach to documentation, ensuring that architectural decisions and standards are recorded clearly and accessibly.
Nice to have
- Hands-on experience with JVM optimization, including tuning, profiling, and advanced troubleshooting of runtime issues.
- Track record of producing high-quality, logical, and detailed technical documentation that supports decision-making and knowledge transfer.
- Demonstrated strength in written and verbal communication, with a proactive approach to coordinating across teams and stakeholders.
- Fluency in both English and Mandarin to facilitate cross-functional collaboration in a multicultural environment.
Practical notes
- Applicants must possess a current right to work in Singapore, as visa sponsorship is not provided for this role.
- Compensation includes a competitive salary, comprehensive healthcare for employees and dependents, meal and wellness allowances, and education subsidies.
- Please apply directly via the official company careers website to ensure accuracy and to access the most current details.
This role is situated within the Service Stability Engineering Team, where you will collaborate with engineers, product managers, and platform specialists. Your contributions will help establish long-term reliability practices that scale as the platform grows. You will be expected to balance deep technical work with strategic planning, ensuring that stability considerations are embedded in every layer of the stack. The position requires comfort with ambiguity and a willingness to tackle difficult problems in a fast-moving financial technology environment. You will work closely with global stakeholders, aligning reliability initiatives with business priorities and regulatory expectations. The successful candidate will be comfortable operating both at the code level and in abstract architectural discussions. This is an opportunity to shape the foundation of critical infrastructure that supports a major digital asset exchange and wallet ecosystem. Your work will have direct implications for customer trust, operational resilience, and the long-term reputation of the platform. The role demands rigor, ownership, and a continuous drive to improve the stability landscape. If you are motivated by challenging technical problems and want to influence architecture at scale, this position offers significant impact and growth.