Member of Technical Staff
Job description
Member of Technical Staff at Basis Ai.
About the role
Basis is building autonomous agents that execute complex, end-to-end workflows for global accounting firms, and this role sits at the heart of that mission. You will serve as the Site Reliability Engineer responsible for designing and managing the infrastructure that ensures our production-grade agent systems run securely, efficiently, and at scale. This is a high-impact position where the lines between infrastructure, product, and machine learning blur and evolve on a frequent basis. You will partner closely with product and engineering teams to translate operational needs into robust technical solutions. The ideal candidate thrives in an environment that demands ownership, deep technical judgment, and a bias for action. You will be expected to mentor peers and elevate the entire team's standards for reliability and operational excellence. Your work will directly influence the stability and performance of systems that handle real economic workflows for some of the world's largest organizations.
Key facts
What you'll do
- Design, deploy, and manage scalable infrastructure that keeps our production systems secure and performant under varying loads.
- Oversee cloud networking, compute, and storage resources while balancing cost constraints and maximizing system availability.
- Build and maintain robust CI/CD pipelines and infrastructure-as-code practices to accelerate and streamline development workflows.
- Lead incident response efforts, including on-call responsibilities, conducting root cause analysis, and documenting post-mortems for continuous improvement.
- Collaborate with product teams to automate capacity planning, system rollouts, and secure data access across environments.
- Provide technical mentorship to engineers and junior staff to improve team standards, tooling, and overall operational culture.
- Partner with machine learning and product teams to ensure that infrastructure supports the rapid iteration and deployment of coding agents.
- Define and enforce monitoring, alerting, and logging strategies to maintain high visibility into system health and performance.
- Evaluate and adopt new infrastructure technologies that align with the long-term vision for autonomous agent workflows.
- Ensure that all systems comply with internal security policies and industry best practices for operational resilience.
- Drive automation initiatives that reduce manual toil and improve the reliability of day-to-day operations.
- Work closely with finance and business stakeholders to translate operational metrics into actionable insights.
- Support the development and maintenance of serverless and containerized workloads in a reliable and scalable manner.
- Contribute to the evolution of the platform by identifying bottlenecks and proposing infrastructure improvements.
Requirements
- Bring 5+ years of professional experience managing production-scale infrastructure in dynamic environments.
- Demonstrate a strong software engineering foundation with proficiency in at least one programming language for automation and scripting.
- Show mastery of cloud architecture principles, security best practices, and database management concepts.
- Have hands-on experience with containerization technologies and modern automation tools used in production systems.
- Prove ability to work on-site in our Flatiron, NYC office 5 days per week to ensure close collaboration and rapid decision-making.
- Exhibit a deep interest in working with coding agents and applying them to solve real-world economic and business problems.
- Display strong ownership mentality and the capacity to operate in a fast-paced, high-responsibility setting.
- Communicate effectively with both technical and non-technical stakeholders to align on goals and constraints.
- Maintain a commitment to continuous learning and adapting to new tools, platforms, and methodologies.
- Understand the importance of reliability, performance, and maintainability in software infrastructure.
Nice to have
- Demonstrate experience with infrastructure-as-code tools such as Terraform or CloudFormation in production settings.
- Show familiarity with observability toolsets including Prometheus, Grafana, OpenTelemetry, BetterStack, or PagerDuty.
- Bring background in managing systems for serverless environments such as Neon, Modal, or similar platforms.
- Have previous work experience in finance, accounting, or other regulated industries.
- Possess contributions to open-source projects related to infrastructure, automation, or observability.
- Show knowledge of modern CI/CD platforms and testing strategies for infrastructure changes.
Practical notes
This role requires on-site presence in the New York Flatiron office five days per week to ensure seamless collaboration across teams. Travel may be required for company events, team offsites, or customer visits as determined by business needs. The position is eligible for full-time employment with comprehensive benefits. Candidates must be authorized to work in the United States without sponsorship for this role. The company offers a structured onboarding process and ongoing learning opportunities to support long-term success.