3 min read

Claude Automates Evals and Hill-Climbing

ClaudeAI evalshill-climbing

Claude documents two commands, build-eval and hillclimb, that turn evaluation design and application tuning into a repeatable workflow. One creates test suites in the codebase; the other iteratively improves configurations against them. Train/test separation matters because it prevents optimization from merely memorizing examples.

Two commands close the evaluation and improvement loop

The key shift is straightforward: Claude brings eval design and application optimization into one repeatable loop. The official Claude blog post Automating eval design and hillclimbing with Claude describes two commands: /claude-api build-eval and /claude-api hillclimb.

The first creates an evaluation suite inside the codebase. Claude’s documentation lists exact-match checks, code-based validation, and LLM-based grading, alongside reference answers and metrics such as accuracy and latency. In this model, an eval is not a separate spreadsheet remembered just before release; it is part of the working project.

The second command iteratively changes the application and measures the result against the tests already created. According to Claude Platform materials, the search can cover system prompts, model selection, effort level, and other API parameters. This is not training model weights; it is automated search across the configuration surrounding the model.

One detail is critical: the data is split into train and test sets, and the final result is checked on a held-out sample. Without that safeguard, hill-climbing could quickly learn to win on a specific set of examples without improving the system’s overall behavior.

As of September 30, 2026, I view this publication as an explanation of an engineering method rather than a dated product release, since the original description gives no announcement date. The underlying idea is already practical at the process-architecture level.

Prompt tuning becomes a search problem

The main consequence is that improving an AI application can move from a series of subjective edits to a measurable cycle. The approach is especially well suited to tasks with reproducible inputs, verifiable outputs, and a stable evaluation function.

I would first inspect three areas: example leakage between train and test, instability in the LLM grader, and whether the metric reflects real quality. If the rubric is weak, hill-climbing will optimize the wrong thing with great discipline. Cost, latency, and failures must also be included in the evaluation; otherwise, the winning configuration may be good on only one metric.

This is not a magic button for improving Claude. It is a way to make experiments reproducible and less dependent on the prompt author’s taste. The hardest component remains unchanged: the system is only as intelligent as its eval is honestly designed.

We previously covered how to measure LLM-as-a-Judge reliability with IRT metrics. That framework helps validate whether automated eval improvements reflect consistent quality rather than noisy scoring.