Pin evaluation to a moment
Every tool call routes through one gateway at the same cutoff time.
MCP GatewayEvaluate agents against time-consistent production data as it actually was. Chalk replays a deployed agent against a dataset of real historical context, with one knowledge cutoff applied to every tool call. Each run pins its dataset, task and scorer versions for consistency, and no eval data leaves your cloud environment.

The same definitions that serve production serve the historical context.
The same cutoff on every tool call, enforced in the MCP gateway.
A run pins the dataset revision, task version and scorer versions.
A scorer function returns a number or result. A judge is a model described in a prompt.
Expand any row into its tool calls, queries, timing and output.
Evals run in your cloud and may use custom vLLMs so no data leaves your VPC.


Every tool call routes through one gateway at the same cutoff time.
MCP GatewayA scorer returns a number, a result, or a list of results.
SCORERS & JUDGESRuns of the same evaluation stay comparable. Any row expands to granular request data.
EVALUATIONS DOCSUpload a dataset or point at a revision, name the task function and the scorers, and run it. Every tool the agent reaches is held to the same cutoff, so the run reflects a real moment rather than today's data.
Evaluations Docs
What we've been up to and where to find us next.
See why Fast Company named us on their list.
Learn how to give agents context for evals and production
How Chalk filled 40+ roles in one quarter without lowering the bar
Read more about our latest product announcement
Talk to an engineer about running agent evals in your own cloud.