Software Engineer, Backend
Job description
About the role
Braintrust operates an observability platform designed for AI agents, enabling organizations to monitor, evaluate, and refine their production models effectively. The company is looking for a backend engineer to build and scale the core infrastructure powering its real-time data platform, which serves both SaaS customers and self-hosted deployments. This position involves tackling complex distributed systems challenges at the intersection of high-throughput data ingestion and low-latency querying. The engineer will play a pivotal role in shaping the reliability and performance characteristics of the platform as adoption grows. Collaboration with product and design counterparts is essential to translate AI-native workflows into robust backend capabilities.
Key facts
What you'll do
- Design and ship features that deliver deep visibility into large language model performance metrics, token consumption, and latency profiles across diverse production workloads.
- Build and operate a unified abstraction layer responsible for normalizing payloads, implementing caching strategies, and proxying requests to a variety of model providers including Anthropic, OpenAI, and Google Gemini.
- Contribute to and maintain open source software development kits that facilitate tracing, logging, and systematic evaluation of LLM interactions for external developers.
- Develop and optimize high-availability data pipelines capable of ingesting, persisting, and querying massive volumes of AI usage telemetry with minimal latency.
- Collaborate closely with product engineering squads to implement sophisticated capabilities such as secure multimodal attachment storage, granular role-based access control systems, and dynamic on-demand function execution environments for agentic workflows.
- Architect database schemas and indexing strategies using Postgres to support complex analytical queries over time-series and event-based data structures.
- Implement infrastructure-as-code patterns using Terraform to manage AWS resources, ensuring reproducible, version-controlled environment provisioning.
- Establish and refine service level objectives (SLOs) and alerting thresholds using observability stacks like Grafana, Prometheus, and Datadog to guarantee platform uptime.
- Conduct code reviews and architectural design sessions to enforce standards for data integrity, security hardening, and horizontal scalability.
- Investigate and resolve production incidents involving distributed tracing gaps, cache invalidation storms, or provider API degradation scenarios.
- Prototype and benchmark emerging database technologies or streaming architectures (e.g., ClickHouse, Kafka) to future-proof the ingestion layer.
- Mentor junior engineers on backend best practices, including concurrency patterns, memory management in Go/Rust, and effective testing methodologies.
Requirements
- Minimum of five years of professional backend engineering experience building and operating production-grade distributed systems.
- Proven track record designing architectures that prioritize horizontal scale, strong data consistency guarantees, and five-nines uptime targets.
- Deep proficiency across a modern polyglot stack including Node.js (TypeScript), Go, Python, and Rust for performance-critical paths.
- Extensive hands-on experience with Postgres for relational modeling, Redis for caching and pub/sub patterns, and AWS primitives (ECS/EKS, Lambda, S3, RDS).
- Expertise in infrastructure automation via Terraform and container orchestration using Docker and Kubernetes or ECS.
- Practical knowledge of observability ecosystems, specifically instrumenting services for metrics, logs, and traces using Datadog, Prometheus, or Grafana.
- Demonstrated ability to decompose ambiguous, high-level product requirements into concrete technical specifications and phased delivery plans.
- Strong cross-functional collaboration skills with a bias for asynchronous communication, comprehensive technical documentation, and knowledge sharing rituals.
Nice to have
- Prior experience building developer-facing tools, API platforms, or observability products for engineering teams.
- Familiarity with the MLOps lifecycle, including model serving infrastructure, feature stores, or experiment tracking systems.
- Contributions to open source projects within the AI/ML ecosystem, particularly around tracing (OpenTelemetry), evaluation frameworks, or agent runtimes.
- Experience optimizing costs for high-volume GPU or inference workloads in cloud environments.
- Background in implementing security compliance frameworks such as SOC 2, HIPAA, or GDPR within SaaS platforms.
- Exposure to stream processing frameworks like Apache Flink, Kafka Streams, or Materialize for real-time analytics.
Skills & tools
- Node.js, Go, Python, Rust
- Postgres, Redis
- AWS, Terraform, Docker
- Grafana, Prometheus, Datadog
- TypeScript, Kubernetes, ECS
- OpenTelemetry, ClickHouse, Kafka
Practical notes
- Comprehensive health benefits package covering medical, dental, and vision insurance for employees and dependents.
- Flexible time off policy encouraging work-life balance without accrual caps or rigid accrual schedules.
- Monthly stipends provided for home office internet connectivity and mobile phone expenses.
- Daily catered lunch, premium snacks, and beverages available at the San Francisco headquarters.
- Total compensation package comprises a competitive base salary benchmarked against top-tier Bay Area startups plus meaningful equity ownership.
- Company observes equal employment opportunity principles, evaluating all candidates without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability.
- On-site presence required five days per week at the San Francisco office; relocation assistance may be discussed for exceptional candidates.
- Visa sponsorship eligibility assessed on a case-by-case basis depending on role criticality and candidate qualifications.