Senior Site Reliability Engineer
Job description
About the role
You will own the reliability and performance of the core financial systems that power DualEntry AI-native ERP, designing and implementing resilient infrastructure on AWS. You will lead the deployment and enhancement of our CI/CD pipeline and Terraform based infrastructure as code, ensuring secure, scalable, and observable operations. You will build and maintain our monitoring and insights pipeline aligned with OpenTelemetry standards to drive proactive system health. You will partner closely with product and engineering teams to eliminate bottlenecks and improve the developer experience. You will champion pragmatic decision making, choosing robust solutions that keep the business moving fast.
Key facts
What you'll do
- Architect and operate cloud infrastructure on AWS to support DualEntry AI-native ERP at scale.
- Define and enforce reliability standards and runbooks to ensure system resilience and uptime.
- Automate infrastructure provisioning and governance using Terraform and related tooling.
- Evolve the CI/CD pipeline to accelerate safe deployments and reduce manual intervention.
- Design and implement observability pipelines with OpenTelemetry to capture metrics, traces, and logs.
- Lead incident response and on-call rotations, driving postmortems and continuous improvement.
- Collaborate with backend engineers to optimize performance, security, and cost across the stack.
- Mentor engineers on best practices for monitoring, debugging, and operating distributed systems.
- Partner with product teams to translate requirements into robust, scalable technical solutions.
- Drive security best practices in infrastructure and deployment workflows to reduce risk.
Requirements
- You have a hardcore work ethic and high agency, taking ownership of complex problems.
- You are familiar with on-call schedules and thrive in fast-paced, production-critical environments.
- You bring at least 5+ years of experience in roles exposed to scalability challenges and improvements.
- You have at least 3+ years of hands-on experience with Terraform and infrastructure as code.
- You have at least 5+ years of experience working with AWS services and cloud architectures.
- You understand backend fundamentals with proficiency in Python.
- You have direct experience with observability tools and OpenTelemetry standards.
- You hold a strong commitment to security, reliability, and operational excellence.
Nice to have
- Experience delivering at scale on projects serving 1M users or similar high-volume systems.
- Hands-on background in backend development and system design.
- A public GitHub profile showcasing impactful open source contributions and infrastructure code.
Practical notes
- This role operates in an in-office environment in New York City with a collaborative, full-time schedule.
- Travel is not required for this position.
- DualEntry sponsors visas and provides relocation support for eligible candidates.
- The start date is aligned with business needs and should be discussed during the hiring process.
Why you will thrive here
You will join a high-trust, high-velocity team that values pragmatism over theory and rapid iteration over bureaucracy. Your feedback will directly shape product direction and infrastructure decisions. You will work alongside outliers who have achieved difficult things in their careers and beyond. You will grow with the company from its early stage and own meaningful impact on a mission-critical platform. Equity, comprehensive benefits, and learning opportunities ensure long-term partnership and development.
Meet the engineering team
Ignacio https://www.linkedin.com/in/ignaciobrasca shaped the baseline of our integration framework from scratch in a single weekend, then spent most weekends coordinating the team while shipping what was never on the roadmap: webhooks, a search engine for millions of transactions, and Puma itself. Pure initiative, compounding.
Marcin https://www.linkedin.com/in/mkuzdowicz took a non-existent tax structure and built it end to end, setting the north star for how the team keeps improving it on top of it; modular and robust, always setting the bar for the rest.
Puhiza https://www.linkedin.com/in/puhiza-doci grew her team from one to full capacity while owning our integration framework end to end; scaling it to support migrations, integrations, and bank connections. The queen of I/O on our platform.
Danylo https://www.linkedin.com/in/godspell33 took ownership of our bank connection engine end to end and rebuilt it from scratch over a weekend.
How you operate
Pragmatic: you like to move forward and make decisions based in reality, not theory. We don't debate if Cassandra has the most theoretical scalability; we use Postgres until it breaks.
Hard working: 'You can work long, hard, or smart, but here you can't choose two out of three' Jeff Bezos.
Curious: you love to learn, are highly curious about new frameworks and solutions to engineering problems.
Fast-moving: you deploy daily, iterate quickly, and never wait for permission.
Benefits & compensation
- Equity valued between $100,000 and $120,000.
- Base salary ranging from $220,000 to $350,000.
- Medical, dental, and vision insurance.
- One Medical membership.
- Mental health support via Talkspace.
- Fertility and family-building support through Kindbody.
- Virtual healthcare with Teladoc Health.
- 401(k) benefits.
- Commuter benefits.
- Learning and development budget for courses, certifications, and language learning.
- Unlimited AI tokens.
- AI skills development and enablement.
- Visa and permanent residence sponsorship.
- Relocation support for employees moving to New York.
- Flexible PTO of 27 days, including public holidays.
- Collaborative, in-office culture with no bureaucracy.
- Snacks, drinks, team events, and more.
Join us to fix the boring, think the hard, and build the impossible, so we become the most robust engineering platform for accounting the world has ever seen.