AI Catchup

Claude Code Can Build Evaluations and Improve Apps Against Held-Out Tests

By 3 min read

Anthropic's Claude API skill gives Claude Code two evaluation workflows: `/claude-api build-eval` creates an evaluation in a codebase, and `/claude-api hillclimb` proposes one change at a time against it. The hillclimb holds out test cases, reverts changes that do not improve the test set, and reports uncertainty against the baseline.

Anthropic's claude-api skill adds two evaluation workflows for Claude Code: /claude-api build-eval creates an evaluation inside a codebase, and /claude-api hillclimb tries to improve an application against that evaluation one change at a time. Anthropic says the hillclimb workflow uses a held-out set to catch overfitting (Anthropic).

The guide is for developers who already have an application and want to measure a concrete change, such as improving performance or lowering cost while holding performance steady. It does not make a benchmark automatically representative of production. The user still chooses the task distribution, reviews the cases and grader, and decides which code or configuration Claude may change (Anthropic).

Build an Evaluation From Real Work

Running /claude-api build-eval interviews the developer, then helps assemble an evaluation set, grader, runner, and baseline. The skill can draw cases from production transcripts, bug reports, support tickets, examples supplied by the developer, or synthetic cases based on the codebase. It shows the inputs for review and pauses at approval points before running the baseline (Anthropic).

The article's design principle is to mirror important production tasks rather than select cases simply because the current model fails them. Anthropic advises using tasks a domain expert can judge consistently, with the criteria stated clearly enough for the grader to check. The baseline also needs headroom: when a top model already scores about 95% or higher, the guide recommends exploring cost or latency rather than trying to improve quality with the same cases (Anthropic).

The skill checks whether the grader gives consistent results on repeated runs and watches for infrastructure problems such as timeouts, API errors, or truncated answers. It reports a baseline score with a confidence interval and recommends reading scored transcripts before trusting the evaluator (Anthropic).

Improve One Change at a Time

Running /claude-api hillclimb asks what to optimize and which changes it may make. The allowed surface can include a system prompt, skills or instruction files, tool descriptions, model choice, effort level, API parameters, or harness code. Claude proposes one patch per round and compares the changed application against the evaluation (Anthropic).

To reduce overfitting, the workflow randomly separates the evaluation into training and test cases. Claude can inspect the training cases, but the held-out test set is kept out of its reach. It reverts a patch when training improves but the test score stays flat, and keeps a patch when both sets improve. If the measured gain is within the evaluation's noise, the guide says not to merge it (Anthropic).

What to Check Before Trusting the Result

An evaluation can still diverge from real production traffic. Anthropic warns that a system can overfit to selected examples or exploit details in the harness instead of improving the application for real users. The guide recommends:

  • Use cases that matter in production, including concrete failures from traffic, tickets, and bug reports (Anthropic).
  • Keep a test set that the optimization loop never sees during tuning (Anthropic).
  • Do not paste failing transcripts into prompts used by the hillclimber, because doing so exposes failures to the optimization loop (Anthropic).
  • Make it structurally difficult for the system being evaluated to access the test answers (Anthropic).
  • Treat a change as meaningful only when its test-set gain exceeds evaluation noise (Anthropic).

These safeguards make the workflow a way to run controlled experiments in an application codebase, not a guarantee that a model change will generalize. Developers should inspect the cases, grader, and transcripts before acting on a reported gain (Anthropic).

Sources

Related coverage

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.