
Software Engineer - Testing Frameworks
Job description
About the role
Baseten is seeking a Software Engineer to own the systems that let every engineer at the company prove their code works before it reaches production. You will define and drive the testing strategy across the organization, setting the standards for correctness, performance, and failure tolerance in a platform where downtime is measured in dollars per second. Your primary ownership will be the full testing stack, from fast and reliable unit test tooling to integration harnesses that spin up realistic environments on demand. You will build resilience and chaos testing capabilities, including fault injection and failure scenarios, to ensure systems degrade gracefully before customers ever see an issue. This is a builder role with org-wide leverage, where you will not write tests for other teams but instead build the frameworks and feedback loops that make writing good tests the path of least resistance. You will attack test suite reliability head-on through flake detection, test impact analysis, and aggressive reduction of end-to-end CI wall time. By mentoring engineers and embedding with product and infrastructure teams, you will raise the collective testing bar and establish testing as a foundational pillar of how Baseten ships AI.
Key facts
What you'll do
- Own Baseten's testing strategy end to end, defining standards, tiers, and tooling that engineering teams build against.
- Build and maintain unit, integration, load, and performance testing frameworks that scale with the velocity of an AI platform.
- Design end-to-end test infrastructure that provisions realistic dependencies on demand, such as ephemeral Kubernetes namespaces and containerized service dependencies.
- Build resilience and chaos testing capabilities, including fault injection, network partition and latency simulation, pod and node failure scenarios, dependency degradation, and automated verification of graceful degradation and recovery.
- Attack test suite reliability head-on with flake detection and quarantine, test impact analysis, test parallelization, and selective execution to reduce end-to-end CI wall time.
- Build observability into the testing layer itself, including coverage reporting, test health dashboards, failure triage tooling, and clear signals about under-tested areas of the codebase.
- Create project templates and shared libraries so new services start with a complete testing setup on day one, reducing friction and enforcing best practices.
- Partner directly with product and infrastructure teams, embedding where needed to understand their hardest testing problems and turning one-off solutions into platform capabilities.
- Mentor engineers across the organization on testing practices through code review, documentation, and internal advocacy to elevate the collective bar.
- Measure and optimize the ROI of testing efforts, balancing thoroughness with speed to keep the test suite fast, deterministic, and trustworthy at scale.
- Extend testing frameworks and their internals, including pytest, Go's testing package, testify, testcontainers, and equivalent tools used by other teams.
- Ensure testing infrastructure runs reliably on real cluster environments, using solid Kubernetes and Docker fundamentals to validate behavior in production-like conditions.
Requirements
- Strong proficiency in Go and/or Python, with hands-on experience building test tooling and libraries used by other engineers.
- Deep experience with testing frameworks and their internals, including pytest, Go's testing package, testify, testcontainers, or equivalents, and the ability to extend them rather than only use them.
- Experience designing integration test infrastructure for distributed systems, including ephemeral environment provisioning and dependency management.
- Practical experience with load and performance testing tools such as k6, Locust, Vegeta, JMeter, or Gatling, and the judgment to interpret results and drive action from them.
- Solid understanding of Kubernetes and Docker fundamentals, with the ability to build test infrastructure that runs on and tests against real cluster environments.
- Advanced understanding of CI/CD, including strategies to keep large test suites fast, deterministic, and trustworthy at scale across many teams and services.
- A track record of measurably improving testing culture on an engineering team, with clear opinions on what to test, what not to test, and where the return on investment actually lies.
- Excitement about building foundational infrastructure and comfort working independently on ambiguous, high-impact technical challenges that affect the entire organization.