CleraVerified Source

Data Scientist — Agent Evaluations & Quality

Onsite · San Francisco, California
Posted August 11, 2026
payroll

Overview

Own the measurement system that determines if our AI executive assistant improves in ambiguous real-world environments. Partner with AI Agent Capabilities engineers to turn hard questions about agent behavior into actionable answers. The role focuses on scaling impact across product surfaces while ensuring quality.

What You'll Do10

  • 1Architect and maintain automated evaluation pipelines measuring agent quality across product surfaces
  • 2Translate product capabilities into explicit pass partial-pass and failure criteria for complex multi-step tasks
  • 3Build representative gold datasets and regression suites covering real workflows edge cases and adversarial scenarios
  • 4Define and track metrics including task success tool-selection accuracy instruction adherence factual consistency latency cost and reliability
  • 5Design deterministic and model-based graders calibrate LLM-as-a-judge systems and monitor grader agreement
  • 6Compare models prompts and implementations using rigorous offline experiments and production evidence
  • 7Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy
  • 8Convert production failures into regression cases and continuously close gaps in evaluation coverage
  • 9Build dashboards and release-quality signals that make results actionable for engineering product and leadership
  • 10Partner with capability engineers to recommend improvements and verify fixes raise quality without unacceptable regressions

Requirements11

  • 14+ years in Applied Data Science or Machine Learning focusing on evaluation systems automated data pipelines or production ML infrastructure
  • 2Proven experience designing automated evaluation frameworks success criteria and regression suites for complex AI/ML or agentic systems
  • 3Production-grade proficiency in Python and SQL with hands-on experience building and maintaining automated analytical pipelines on large datasets
  • 4Core requirement applying statistical and experimental methods significance testing variance analysis sampling to evaluate non-deterministic AI/ML systems
  • 5Experience developing labeled datasets annotation guidelines and quality-control processes for ground-truth data in dynamic product environments
  • 6Deep understanding of LLM agent behaviors including tool use multi-step execution retrieval and practical failure modes
  • 7Ability to analyze model traces tool calls and outputs to identify root causes of failures across model prompt tool and data layers
  • 8Experience using production telemetry and observability data to monitor system quality build dashboards and analyze real-world user outcomes
  • 9Nice to have prior hands-on experience with LLM-as-a-judge systems model-based grading or AI benchmarking platforms
  • 10Experience shipping or operating production ML products agentic systems or customer-facing consumer software
  • 11Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems

Salary Insight

Salary not disclosed in listing

Location

Typeonsite
LocationSan Francisco, California

Required Skills

PythonSQLStatistical methodsSignificance testingVariance analysisSamplingLabeled datasetsAnnotation guidelinesLLM agent behaviorsModel trace analysisProduction telemetryDashboardingLLM-as-a-judge systemsModel‑based grading
Share:

Similar open positions

Explore active roles that match your skills and interests.

Clera

10h agoRemotepayroll

Data Scientist, Agent Evaluations & Quality

You will own the evaluation and quality measurement systems for Clera's AI agents, which handle email, calendar, and browser automation across diverse business software. As part of a 25-person team, you will define success criteria, build datasets, and design grading pipelines that drive engineering decisions. Working with Python, SQL, and LLM-as-a-judge frameworks, you will calibrate model-based graders and measure false positives, variance, and agreement. This role bridges applied data science and production quality engineering, offering direct impact on agent reliability and user trust.

Competitive salary
pythonsqlllm+2 more
Wise Skulls Corp.

Wise Skulls Corp.

6h agoAustin, Texaspayroll

AI Engineer (LLM Agents & Data Engineering)

Lead design and delivery of AI solutions that scale across multiple platforms. Own the end-to-end pipeline from concept to production while driving innovation in large language models. This role shapes how our systems learn and adapt.

Competitive salary
PythonLLMsPrompt Engineering+8 more
ADUS-Adobe Inc.

ADUS-Adobe Inc.

15h agoSan Jose, Californiapayroll

Senior Software Engineer, Agentic Systems & LLM

You will design and build the agent ecosystems and orchestration systems that move Generative AI research into production across Adobe's flagship creative products, impacting millions of creators using Photoshop, Lightroom, and Illustrator. Partner with Adobe Research and platform teams to define the standards and infrastructure for agent-generated code. This hands-on senior IC role defines a new engineering discipline: building systems that build software. You will own the evaluation frameworks and production pipelines that ensure agent quality, reliability, and safety at scale.

152K–265K
PythonPyTorchTensorFlow+9 more

Everforth, Cybercoders

10h agoRemotepayroll

Senior AI/ML Engineer, LLM & Agent Orchestration

You will design, build, and operate production AI systems that power the intelligence layer for enterprise portfolio management. You'll join a product team shipping LLM-powered features used by 30,000+ customers, including Amazon, Disney, and PayPal. The stack includes Python, Kotlin, TypeScript, and AWS Bedrock with Claude. You'll own the AI components from signal detection to agent orchestration, working alongside domain engineers and product managers. This role stands out for its focus on evaluation-driven development and production discipline rather than research silos.

220K–300K
PythonKotlinTypeScript+3 more

Ambiencehealthcare

6h agoSan Francisco, Californiapayroll

Senior Machine Learning Engineer Ambience Healthcare

Lead development of trustworthy AI evaluation systems and production model behavior at Ambience. Own end-to-end AI systems across models data and infrastructure. Drive measurable improvements in LLM and agentic platforms.

Competitive salary
trustworthy ai evaluation systemsproduction model behavioragentic ai systems+2 more

David Joseph & Company

10h agoSan Francisco, Californiapayroll

Applied AI Engineer, Agent Harness & LLM Systems

You will own the core agent harness that turns raw model capability into dependable product for real users. You'll build the execution loop, tool-use strategies, and context construction, and shape agent behavior across real customer workflows. You'll work with Python or TypeScript, modern AI tooling, and sandboxed microVM execution. The role sits on a small founding team at a seed-stage startup backed by strong operators, with design partners already live. You'll define the boundary between runtime model decisions and deterministic, compiled code.

200K–300K
PythonTypeScriptAgent frameworks+3 more