
Data Scientist — Agent Evaluations & Quality
About the Role
This company is building an AI executive assistant that operates across email, calendars, meetings, and business software. As a Data Scientist — Agent Evaluations & Quality, you will own the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments. You'll partner directly with AI Agent Capabilities engineers to generate the evidence that shapes product decisions, model choices, and release quality.
This is a high-ownership, deeply technical role at the intersection of applied data science, LLM evaluation, and product quality — ideal for someone who thrives on turning hard, open-ended quality questions into rigorous, actionable answers.
What You'll Do
Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.
Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.
Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.
Define meaningful metrics — task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.
Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement.
Compare models, prompts, and implementations using rigorous offline experiments and production evidence.
Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.
Turn production failures into regression cases and continuously close gaps in evaluation coverage.
Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.
Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions.
What We're Looking For
Required
4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.
Experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.
Production-grade proficiency in Python and SQL, with experience building and maintaining automated analytical pipelines on large datasets.
Applied statistical and experimental skills: significance testing, variance analysis, and sampling to evaluate non-deterministic AI/ML systems.
Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments.
Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, and practical failure modes.
Demonstrated ability to analyze model traces, tool calls, and outputs to identify root causes across model, prompt, tool, and data layers.
Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes.
Nice to Have
Hands-on experience with LLM-as-a-judge systems, model-based grading, or AI benchmarking platforms.
Experience shipping or operating production ML products, agentic systems, or customer-facing consumer software.
Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems.
What makes you a great fit
You're product-oriented — you prioritize metrics tied to real user outcomes, not just convenient measurements.
You drive ambiguous quality questions from evaluation design all the way into product decisions.
You write maintainable, production-quality code — not just ad-hoc notebooks.
You collaborate naturally with engineers and are comfortable digging into traces and system internals.
Location
This role is on-site. Visa sponsorship is not available for this position.
Compensation & Benefits
Compensation details were not provided for this listing. A competitive package commensurate with experience is expected at this stage of company growth.
Find more English Speaking Jobs in United Kingdom on Arbeitnow
Skills
- Engineering
Similar roles
Data Analyst - Full-Time/ Permanent
TECH WIRE IT PTY LTD·Australia
About the company Tech Wire IT Pty Ltd is an Australian information technology company engaged in the provision of software development, IT consulting, system support, and digital technology solutions. The company is committed to delivering innovative, reliable, and high-quality technology services that help businesses improve operational efficiency and achieve their strategic objectives. Role overview The Data Analyst is responsible for…
- Full-time
- AUD 90k–AUD 110k / yr
- Skilled Independent 189 — PR, no sponsor
Sr Manager, Field Enablement, North America
Rubrik Job Board·Worldwide
Sr. Manager of Field Enablement, North America As a Sr. Manager of Field Enablement for North America at Rubrik, you occupy a high-impact leadership role at the intersection of strategy and field execution. Rubrik operates in a hyper-growth environment, and this role is specifically designed to bridge the gap between global corporate strategy and the regional nuances of the North…
- Full-time
Analytics Engineer, GTM
Openai·San Francisco, United States
About the Role As a member of the Data team within the Go-to-Market organization, you will help build a data-driven culture, improve decision-making, and advance strategic initiatives through analytics. This is a full-stack data role spanning data modeling, metric definition, visualization, analysis, and self-service tooling. You will build trusted, scalable data sources and products that give the business reliable, actionable…
- Full-time
Senior Privacy Specialist
Davy·Dublin, Ireland
Abous Us At Davy, it’s the unique talents of all our people that have been the foundations of our success for 100 years. As we continue to grow, so do you. Because you are not just part of our team – you are a key player in shaping our future. At Davy, you are the difference. Established in 1926, the…
- Hybrid
- Full-time
- Critical Skills — 9mo tied, then free; Stamp 4 @21mo
Junior Data Scientist
Livescore9·London, Canada
<p><strong>Soho, London <br></strong><strong><em>Hybrid working: 3 days in the office Tues - Thurs</em><em><br><br></em></strong></p> <p><strong>The Role</strong></p> <p>We are looking for a brilliant and motivated Junior Data Scientist to join our Marketing Analytics team. In this role, you will help us measure, optimize, and supercharge our marketing investments across our core brands, with a strong focus on our mathematical and econometric modelling frameworks. </p> <p>You…
- On-site
- Full-time
- Express Entry — PR day one, no employer

Data Scientist
Weetabix·Burton Latimer, United Kingdom
DATA SCIENTIST X4 - FTC At Weetabix, we believe that diverse teams drive better ideas, stronger decisions, and a more inclusive workplace for everyone. We’re committed to building an organisation where people from all walks of life feel they belong—where different voices, experiences, and backgrounds are valued and respected. Closing date: 20th August 2026 Interview process: 1st Stage Interview followed…
- Hybrid
- Full-time
SAP PP/QM Consultant (Senior)
Clera·Worldwide
ABOUT THE ROLE We are seeking an experienced Senior SAP PP/QM Consultant for a long-term, full-time onsite contract engagement in Alpharetta, GA. This is a Day 1 onsite role requiring candidates who are available and able to work on-site from the start. The ideal candidate brings 12–15+ years of hands-on SAP functional consulting experience with deep expertise in Production Planning…
- On-site
- Full-time
Community Operations Lead
Clera·Worldwide
ABOUT THE ROLE We're an early-stage consumer tech / AI company looking for a Community Operations Lead to build and own our community presence from the ground up. You'll be the central figure across the platforms where our early users live — driving engagement, shaping community culture, and turning our most passionate users into true product evangelists. This is a…
- On-site
- Full-time