Data Scientist
Job description
About the role
Arena Intelligence is a data-centric organization dedicated to assessing real-world AI performance. Our origin lies in UC Berkeley's SkyLab, and our mandate is to measure, refine, and understand artificial intelligence as it functions in practical settings. Each month, millions of users evaluate frontier systems through the tasks they actually perform, from agentic coding to creative production and professional workflows. The insights they generate fuel the most transparent and rigorous evaluations in the field, guiding leading labs, enterprises, and independent researchers. We move beyond static rankings to dissect the human experience, ensuring models evolve to serve actual work.
Our team blends researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We operate with speed, prioritize craftsmanship and curiosity, and measure impact rather than status. Our culture demands excellence, energy, and focus, and it welcomes thoughtful people from all backgrounds. This role centers on experimentation, causal inference, and retention analytics. You will translate complex user behavior into clear, data-backed decisions that enhance engagement. You will manage large-scale evaluation and build measurement frameworks that withstand scrutiny. Expertise with PySpark is a significant advantage for handling our data volumes.
Key facts
What you'll do
- Drive experimentation strategy by designing and executing A/B tests, multi-armed bandits, and quasi-experimental studies to quantify product impact.
- Apply causal inference methods such as difference-in-differences, propensity score matching, synthetic control, and regression discontinuity to assess effects in non-randomized environments.
- Define hypotheses in collaboration with product, engineering, and marketing partners and clarify success metrics and statistical power requirements.
- Construct and refine retention measurement frameworks tracking DAU, WAU, MAU, rolling retention, and N-day retention to monitor long-term engagement.
- Perform cohort analysis, survival analysis, and churn prediction to identify drivers of sustained user engagement.
- Build data pipelines using PySpark, SQL, and big data technologies to process high-volume information reliably and efficiently.
- Create dashboards in tools such as Tableau, Looker, or Metabase to communicate experiment results and retention trends to stakeholders.
- Conduct segmentations, funnel reviews, and predictive modeling to pinpoint where improvements will most strongly affect growth and retention.
- Partner with growth teams to translate insights into action, enhancing onboarding, engagement loops, and monetization strategies.
- Uphold rigorous analytical standards including bias control, multiple testing adjustments, and precise confidence intervals to ensure decision integrity.
Requirements
- Bring a minimum of three years of professional experience in data science, analytics, or experimentation roles.
- Demonstrate strong statistical knowledge including hypothesis testing, Bayesian methods, and experimental design principles.
- Apply SQL and Python libraries such as Pandas, NumPy, SciPy, StatsModels, and Scikit-learn to analyze complex datasets.
- Utilize experimentation platforms such as Optimizely, Statsig, Eppo, or custom in-house systems to run and evaluate studies.
- Define and analyze retention metrics confidently, showing a track record of measuring and improving engagement.
- Work comfortably with big data tools including PySpark, Hadoop, or similar frameworks to process and analyze high-volume information.
- Communicate analytical findings clearly to both technical and non-technical audiences through reports and presentations.
- Collaborate effectively with cross-functional teams including product managers, engineers, and marketers to align analysis with business goals.
Nice to have
- Apply advanced PySpark techniques for large-scale data processing and optimization.
- Build time-series forecasting models to anticipate user behavior and engagement trends.
- Perform survival analysis and uplift modeling to deepen causal understanding of user journeys.
- Develop clustering and recommendation systems to enhance personalization and retention.
- Create data visualizations using Tableau, Looker, Plotly, Matplotlib, and Seaborn to strengthen storytelling.
- Contribute domain knowledge in growth, product, or marketing analytics to refine measurement strategies.
Practical notes
- Location is specified as the Bay Area.
- Engagement type is full_time.
- No compensation details are provided in the source material.
- No specific hours, travel requirements, visa sponsorship details, or application deadlines are stated in the source.
The role requires a data scientist who thrives in a fast-paced, measurement-focused environment and is capable of turning complex user behavior into actionable insights. You will own the end-to-end analytical lifecycle, from hypothesis generation and experimental design to insight communication and framework maintenance. Your work will directly influence product decisions and strategic direction for Arena Intelligence, helping to ensure that evaluations of AI systems remain transparent, rigorous, and aligned with real-world needs. If you are passionate about causality, retention analytics, and scalable data processing, this position offers the opportunity to build measurement infrastructure used by leading labs and researchers. Success in this role depends on your ability to combine statistical rigor with practical impact, using data to tell clear stories about how AI performs in the hands of real users.