Agentic Coding Articles
12 articles across AI Catchup's news, guides, tutorials, and comparisons.
All Agentic Coding articles
Claude Code Can Build Evaluations and Improve Apps Against Held-Out Tests
Anthropic's Claude API skill gives Claude Code two evaluation workflows: `/claude-api build-eval` creates an evaluation in a codebase, and `/claude-api hillclimb` proposes one change at a time against it. The hillclimb holds out test cases, reverts changes that do not improve the test set, and reports uncertainty against the baseline.
Claude Sonnet 5.5 Launches With Lower Task Costs and Cursor Support
Anthropic introduced Claude Sonnet 5.5 on September 28 with model ID `claude-sonnet-5-5`. It keeps Sonnet 5's $2 input and $10 output price per million tokens; Anthropic reports output generation more than 30% faster and up to 30% lower cost per task. Cursor lists the model in Settings > Models. Cache reads dropped to $0.10 per million tokens on October 7.
Claude Opus 5.5: Fable 5.1-Level Scores at $4 and $20, With Benchmarks, Effort Costs, and Migration Guide
Claude Opus 5.5, released September 22, 2026, beats Claude Fable 5.1 on every benchmark Anthropic published and costs $4 and $20 per million tokens, 60% less than Fable 5.1. It scores 66.4% on Terminal-Bench 4.0 against 57.9% for GPT-6 Astra. Thinking is always on, the default effort is medium, and four API changes can break Opus 5 code.
OpenAI Welcomes the Git AI Team Behind an Open-Source Agent Attribution Tool
OpenAI says Aidan Cunniffe and Sasha Varlamov from the Git AI team have joined OpenAI. Their open-source Git extension tracks AI-generated code at line level, linking it to the agent, model, and prompts that produced it, with commands for attribution stats and AI-aware blame.
Cursor Projects Give a Coordinator Agent a Long-Lived Body of Work
Cursor's Projects beta moves beyond one chat at a time: a coordinator agent keeps context over months, delegates implementation and testing to subagents, shares research and artifacts across cloud and local machines, and can watch PRs, Slack, or a schedule for recurring work.
OpenAI Opens Its Codex-Powered Agents API to All Developers in Public Beta
OpenAI's Agents API is now in public beta for all developers. The managed service uses the Codex harness to run long-lived cloud agents, coordinate parallel subagents, compact context across long sessions, and work in OpenAI-hosted, VPC, or partner-provided sandboxes.
Cursor Cloud Agents Can Now Run on Self-Hosted Machines
Cursor Cloud Agents can now execute on machines that a team manages, including dynamically scaling pools inside its network. The execution environment moves to the team's infrastructure while Cursor continues to handle the agent loop, inference, and planning.
Cursor Cloud Agents Add Event Triggers, Long-Lived Goals, and Isolated Subagents
Cursor says Cloud Agents can now pick up work from events, keep working toward a long-lived goal, monitor pull requests, watch Slack threads, run scheduled tasks, and launch subagents in isolated virtual machines. The same update adds skill-based Custom Modes and less disruptive steering while an agent is working.
Cursor Cloud Agents Start Up to 3x Faster With Builds
Cursor says Cloud Agents can start up to 3x faster with Builds, ready-to-use copies of development environments prepared in the background. Builds keep agents on the latest successful environment, expose logs and version history in the dashboard, and are now the default path for every Cloud Agent environment.
Grok 4.6 Launched in Cursor With a One-Week 2x Usage Window, Now Ended
SpaceXAI released Grok 4.6 on August 12, 2026, focused on long-running agents and visual work. Its first-week 2x included-usage offer in Cursor and Grok Build has ended, and Grok 4.7 succeeded it on September 21 at the same price. Grok 4.6 API pricing starts at $2 per million input tokens and $6 per million output tokens.
Claude Code Made Auto Mode the Default on August 14, and Now on Every Plan
Auto mode is the built-in starting permission mode for interactive Claude Code terminal and VS Code sessions on every plan and provider with v2.1.283 or later. It became the default for Pro, Max, and Team on August 14, 2026. A classifier reviews each tool call, and pinned or managed defaults still win.
Cursor Composer 2.5: Better Long-Running Agent Work, Standard vs Fast Pricing, and Composer 2 Retired
Composer 2.5 is Cursor's own agentic coding model and, as of October 2026, the only Composer model Cursor lists: Composer 2 is retired, and SDK requests for it reroute to Composer 2.5. Cursor says 2.5 is better at long-running tasks and complex instructions. Standard costs $0.50/$2.50 per million tokens; Fast, the default, costs $3/$15.