Agent Evals

Accurate evals built on real data, for agents and models

Evaluate agents against time-consistent production data as it actually was. Chalk replays a deployed agent against a dataset of real historical context, with one knowledge cutoff applied to every tool call. Each run pins its dataset, task and scorer versions for consistency, and no eval data leaves your cloud environment.

hero gradient
Whatnot logo
Socure logo
Sunrun logo
Grindr logo
Turo logo
MoneyLion logo
Melio logo
Mission Lane logo
Medely logo
Iwoca logo
Nowsta logo
Apartment List logo
Pipe logo

Explore Chalk Evaluations

Production data is the eval set

The same definitions that serve production serve the historical context.

One cutoff, whole trajectory

The same cutoff on every tool call, enforced in the MCP gateway.

Immutable definitions

A run pins the dataset revision, task version and scorer versions.

Scorers & judges

A scorer function returns a number or result. A judge is a model described in a prompt.

Read the whole run

Expand any row into its tool calls, queries, timing and output.

Keep data in your cloud

Evals run in your cloud and may use custom vLLMs so no data leaves your VPC.

company logo

Compared to some problems, we make relatively few, very high-value decisions. Throughput isn't the problem. Trust is.

Rich Pearce
Rich PearceEngineering Director, iwoca

How Agent Evals work

Pin evaluation to a moment

Every tool call routes through one gateway at the same cutoff time.

MCP Gateway

Score with functions or judges

A scorer returns a number, a result, or a list of results.

SCORERS & JUDGES

Compare runs, then open one

Runs of the same evaluation stay comparable. Any row expands to granular request data.

EVALUATIONS DOCS

Replay the world as it was

Upload a dataset or point at a revision, name the task function and the scorers, and run it. Every tool the agent reaches is held to the same cutoff, so the run reflects a real moment rather than today's data.

Evaluations Docs
product section resource

Keep up with Chalk

What we've been up to and where to find us next.

Know which agent change is shippable

Talk to an engineer about running agent evals in your own cloud.