Site Reliability Engineer
Job description
About the role
Arena Intelligence is seeking an engineer to develop the foundational infrastructure powering our AI evaluation platform. You will manage the gateways, runtimes, and serving layers that facilitate large-scale model assessment for users worldwide. This role requires a hands-on operator who builds the systems that ensure reliability, scale, and performance are non-negotiable. You will own the critical infrastructure that connects complex AI workloads with the demands of production environments. The ideal candidate thrives on solving ambiguous problems that sit at the intersection of distributed systems and machine learning. You will be responsible for the uptime, scalability, and efficiency of the core platforms supporting our customers. This position is for a self-starter who is not afraid to dive deep into complex systems to enable the company's rapid growth.
Key facts
What you'll do
- Architect and maintain low-latency APIs that power model leaderboards and evaluation arenas for a global user base.
- Design and implement enterprise-grade features such as audit logging, rate limiting, authentication, usage metering, and SOC 2 compliance controls.
- Establish comprehensive observability by deploying distributed tracing, real-time dashboards, and token-level usage monitoring to ensure system health.
- Collaborate closely with researchers to translate experimental data and benchmarks into robust, production-ready features and integrations.
- Build and manage the serving layers that handle the dynamic nature of large language model inference and evaluation workloads.
- Optimize the infrastructure for high throughput and efficient resource utilization to meet the demands of scaling AI workloads.
- Define and enforce reliability standards and best practices across the engineering organization to prevent outages and degradation.
- Work with cross-functional teams to diagnose complex system issues and implement resilient solutions that improve the overall architecture.
- Ensure the platform can handle variable loads while maintaining strict security and compliance requirements for enterprise customers.
- Contribute to the technical roadmap by proposing improvements to the infrastructure that enhance scalability and developer experience.
Requirements
- Possess 6+ years of backend engineering experience with a strong focus on distributed systems or developer platform technologies.
- Demonstrate proficiency in Go or Rust, with a proven track record of building high-throughput proxy or gateway systems that handle significant traffic.
- Have practical knowledge of major LLM provider APIs and their operational nuances, including handling streaming responses, token management, and strict rate limits.
- Show expertise in modern cloud infrastructure, specifically AWS or GCP, along with Kubernetes orchestration, Terraform for infrastructure as code, and data stores like Postgres and Redis.
- Exhibit a product-focused mindset that prioritizes exceptional developer experience and clear, logical problem-solving in complex scenarios.
- Display the ability to thrive in a fast-paced startup environment where priorities shift rapidly and adaptability is essential.
- Bring a strong sense of ownership to troubleshoot and resolve critical production incidents with minimal disruption to services.
- Commit to writing clean, maintainable code and documenting systems effectively to support long-term scalability and team collaboration.
Nice to have
- Hands-on experience with API gateways or proxies such as Envoy, Kong, Tyk, or Bifrost to manage traffic and security.
- A background in ML infrastructure, model serving pipelines, or evaluation frameworks that bridge the gap between research and production.
- Familiarity with enterprise requirements including Single Sign-On (SSO), Role-Based Access Control (RBAC), and multi-tenancy architectures.
- Experience with billing systems and platforms like Stripe, Orb, or Metronome to support monetization and usage tracking.
- Knowledge of the current AI stack and tools such as LiteLLM, vLLM, or LangChain to optimize and debug model interactions.
Skills & tools
Go, Rust, AWS, GCP, Kubernetes, Terraform, Postgres, Redis, LLM APIs, Distributed Systems.
Practical notes
We offer competitive compensation and equity packages tailored to your location. Benefits include medical, dental, and vision coverage. Arena is an equal opportunity employer committed to a diverse and inclusive workplace.