Staff Software Engineer, Product
Job description
About the role
Arena Intelligence is a research-driven company focused on evaluating AI performance in real-world conditions rather than controlled benchmarks. Founded by researchers from UC Berkeley's SkyLab, the organization measures how frontier systems behave outside labs to understand their actual impact on work. Each month, tens of millions of users evaluate frontier models through the lens of their daily tasks, generating the human-centric data that defines the company's methodology. This role owns specific product areas end-to-end, from discovering unaddressed opportunities to shipping solutions that move measurable outcomes. The hire acts as the primary architect, writing code and designing systems that directly shape how the platform evaluates agentic coding, creative generation, and professional productivity. They establish the technical standards that allow the most transparent and rigorous evaluation methodology in the AI sector to scale. Success is defined by creating durable impact that improves how models serve the work people genuinely perform.
Key facts
What you'll do
- Assume end-to-end ownership of product areas, discovering opportunities absent from roadmaps and shipping solutions that generate measurable impact.
- Design complete system architectures, defining data models, API contracts, frontend architecture, backend services, infrastructure, and deployment strategies.
- Make high-stakes technical decisions amid ambiguity, choosing between build versus buy and monolith versus service while defining boundaries between prototyping and hardening.
- Translate research services from ML researchers into reliable production systems capable of handling real-world evaluation workloads.
- Drive post-launch iteration by measuring results, correcting course, and ensuring sustained impact rather than treating version one as a final deliverable.
- Elevate the technical and product quality ceiling for the surrounding team by introducing new practices, unblocking colleagues, and clarifying ambiguous situations.
- Navigate cross-functional collaboration with product, design, research, and leadership to synchronize on priorities aligned with real-world workflows.
- Champion the domain of AI evaluation, articulating its significance and guiding the product's evolutionary path through clear communication.
- Write high-quality code that directly powers evaluations of models in areas such as agentic coding, creative generation, and professional productivity.
- Analyze user behavior and preference data to refine evaluation methodologies and improve how models are assessed across diverse workflows.
- Ensure system designs support transparency, rigor, and human-centered evaluation by considering the entire stack from data flow to user experience.
- Balance rapid experimentation with the need for robust, hardened production systems that stakeholders can trust.
- Maintain a deep understanding of performance constraints such as latency, throughput, and cost when designing evaluation platforms.
- Serve as the definitive expert for specific product domains, providing direction and mentorship to other engineers on technical and product decisions.
Requirements
- Bring more than eight years of software engineering experience centered on product development in a professional setting.
- Demonstrate deep experience in building web applications, requiring mastery of data models, APIs, frontend architecture, and deployment strategies for complex systems.
- Maintain a track record of repeated, attributable impact across multiple projects and roles, with the ability to quantify outcomes such as revenue, engagement, or efficiency that persisted beyond the duration of employment.
- Exhibit personal technical depth, engaging with real performance challenges such as slow queries, N+1 issues, caching strategies, and transaction isolation while discussing decisions with specificity.
- Show product judgment through opportunity identification, evidence-based validation, securing buy-in, shipping strategic bets, and recognizing when to terminate initiatives that do not meet expectations.
- Communicate with clarity and persuasion to establish alignment across engineering, product, and leadership teams, creating understanding for stakeholders beyond the immediate technical context.
- Demonstrate genuine conviction regarding AI evaluation and Arena's mission, with the ability to concretely explain the importance of the domain and the future direction of the product.
- Operate effectively within a fast-paced, truth-seeking environment that values curiosity, speed, and craftsmanship above corporate hierarchy.
Nice to have
- Production-level experience with AI or LLM systems, including inference pipelines, evaluation workflows, model integration, or AI-powered product features.
- Familiarity with the company's technology stack, which includes NextJS, React, TypeScript, Tailwind, ShadCN, HonoJS, Postgres, and Vitest.
- Experience utilizing Supabase or Vercel's AI SDK for building scalable AI applications.
- A history of raising the performance bar on a team through the introduction of practices, tools, or standards that were subsequently adopted by others to improve quality and velocity.
Practical notes
This role requires consistent presence in the Bay Area office. Travel may be necessary. The company does not offer visa sponsorship. Standard full-time hours apply without expectations of guaranteed overtime beyond normal workloads.