1 article across AI Catchup's news, guides, tutorials, and comparisons.
Anthropic's Claude API skill gives Claude Code two evaluation workflows: `/claude-api build-eval` creates an evaluation in a codebase, and `/claude-api hillclimb` proposes one change at a time against it. The hillclimb holds out test cases, reverts changes that do not improve the test set, and reports uncertainty against the baseline.
Tools, practices, and what matters, in your inbox every week.