Data Scientist, Agent
Job description
About the role
You are responsible for how Lovable measures and improves its AI agent. You design the evaluation systems and experiments that reveal whether a change makes the agent better or worse, and you translate agent telemetry into the fixes that raise success rates and reduce errors. At Lovable, data scientists are not isolated model builders; they work closely with the product team, running experiments continuously to understand how intelligence changes user behavior and product dynamics. Lovable is a small, talent-dense company based in Stockholm, building a generation-defining platform that lets anyone create software in any language.
Key facts
What you'll do
Define and own the metrics for agent quality, including success rates, completion rates, error rates, and the behaviors that drive those outcomes. Build the evaluation systems and experiment framework that determine whether an agent change ships, such as an A/B-tested rollout that catches a change increasing errors before it reaches all users. Turn agent traces and telemetry into concrete fixes by working directly with the agent engineering team. Build the tooling and agents that produce these evaluations continuously as the agent evolves. Set the standard for how the company judges agent behavior in situations where there is no answer key to check against. Collaborate with the product team to understand how intelligence changes user behavior and product dynamics. Design A/B tests for agent changes where outcomes are noisy and hard to measure cleanly. Create systems that generate insight on an ongoing basis rather than producing one-off analyses. Identify what "good" agent behavior looks like and develop ways to measure it even without a clear ground truth. Work closely with engineers building the agent to turn data into actionable improvements. Own the feedback loop between agent performance data and the engineering changes that follow.
Requirements
Strong SQL and Python skills applied to real-world data problems. Experience or a strong interest in LLM evaluation and observability, including building evals, scoring outputs, tracing agent behavior, and catching regressions. Applied statistics knowledge and comfort designing experiments, especially A/B tests for agent changes where outcomes are noisy. An instinct for what good agent behavior looks like, including success, error rates, and task completion. The ability to set measurement standards for agent behavior when there is no clean answer key available. An entrepreneurial mindset that thrives in ambiguity and works closely with the engineers building the agent. Experience building systems and agents that produce insight continuously rather than through one-off analyses. A background in data science with a focus on making AI agents measurably better rather than just reporting on them.
Nice to have
Familiarity with Braintrust for LLM evaluation and observability. Experience with OpenTelemetry tracing for agent behavior analysis. Exposure to BigQuery and PubSub for warehouse and event data. Comfort using Hex and Lovable Apps for analytics and product insights. Prior experience with growth testing and experimentation frameworks. Knowledge of multiple LLM providers and their APIs.
Skills & tools
SQL for querying and analyzing agent telemetry at scale. Python for building evaluation pipelines and analysis tooling. Braintrust for LLM evaluation and observability. OTEL tracing for capturing and analyzing agent behavior across requests. BigQuery as the primary data warehouse. PubSub for event streaming and real-time data pipelines. Hex for analytics and product insights. Lovable Apps for internal product analytics. A/B and growth testing methodologies for experimentation. GCP cloud infrastructure for deploying and running evaluation systems.
Practical notes
The application process starts with filling in a short form and jumping on an intro call with the recruiting team. This is followed by a call with the hiring manager, a take-home case study, and a Most Impressive Project session. Candidates then go through cross-functional interviews with the people they would work with, and finish with a final conversation with leadership. All applications must be submitted in English, as it is the company language. Lovable treats all candidates equally and is an equal opportunity employer. Applications should be submitted through the company careers portal. The team is based in Stockholm and values extreme ownership, high velocity, and low-ego collaboration.