AI Agents Articles
24 articles across AI Catchup's news, guides, tutorials, and comparisons.
All AI Agents articles
Prompting Claude Opus 5.5: Delete 'Think Carefully', Name the Finish Line, Start at Medium
Prompt Claude Opus 5.5 by handing over the whole task with a clear finish line, deleting “think carefully” lines, and starting at medium effort. Thinking is always on, so effort is your depth control. Add a CLAUDE.md rule that names when to stop, and set effort explicitly instead of carrying Opus 5 settings over.
Prompting Claude Sonnet 5.5: Effort, between_tools, and the Five Breaking Changes
Run Claude Sonnet 5.5 at high effort for general API work, medium for well-specified agentic coding, and medium or low for chat. Thinking is on by default; turn it off with between_tools, not disabled. Five API changes break Sonnet 5 code, and cache reads now cost $0.10 per million tokens.
Prompting GPT-6 Astra and GPT-6.1 Sol: Effort, Migration, and Prompts That Keep It Working
GPT-6 Astra follows instructions more closely than earlier OpenAI models and asks more questions, so prompting it is mostly removing old scaffolding: grant permission to finish, trim skills and AGENTS.md, and stop ordering tests. In the API, replace none effort with low, move tool calls to Responses, and switch caching to prompt_cache_options.ttl.
Scheduled Claude Agents: Pick the Right Surface, Then Stop Them Failing Silently
Run repository work on Claude Code routines and production automations on Claude Managed Agents scheduled deployments, which add minute-level cron and a dollar cap per run. Whichever surface you pick, apply Anthropic's six rules: bookmarks, unreadable-not-quiet, re-check before posting, confirm then record, read-only preferences, least privilege.
OpenAI Launches dots: Always-On Agents in ChatGPT With Their Own Cloud Computer
OpenAI's dots are always-on agents in ChatGPT, powered by GPT-6 Astra, that keep working between conversations using their own cloud computer and the apps you connect. They are rolling out gradually to eligible Pro users outside the EEA, Switzerland, and the UK and to Business Premium, and usage won't count toward eligible plan allowances for the next month.
OpenAI Developers Names 10 WebMCP Challenge Winners
OpenAI Developers announced 10 WebMCP Challenge winners. The projects show websites exposing structured tools that agents can use alongside people, from editable floor plans and wedding seating to 3D scans, notebooks, and fantasy maps.
Cursor Rollouts Watches Deployments for Regressions Before Users See Them
Cursor introduced Rollouts, a Teams and Enterprise bot that follows a change from pull request to production. It writes an editable monitoring plan, checks deploys against logs, metrics, and traces, and reports verified health, regressions, or inconclusive results. It does not merge or roll back on its own.
Claude Code Adds Configurable AGENTS.md Support
Claude Code 2.1.277 adds AGENTS.md support through a built-in mod. By default Claude Code reads AGENTS.md when a project has no CLAUDE.md or CLAUDE.local.md, and /config changes the behavior. Since 2.1.281 it also works on Amazon Bedrock, Google Vertex AI, Microsoft Foundry, LLM gateways, and with telemetry disabled.
Claude Code Projects Turn One Conversation Into Parallel Cloud Sessions
Anthropic's redesigned Projects beta gives Claude Code one coordinating conversation that delegates work to parallel threads. Threads are usually cloud sessions, which Anthropic's docs now list as available, no longer in research preview, on Pro, Max, and Team and for eligible Enterprise seats. A project can also run a thread on your own machine through Remote Control.
OpenAI Shares a Defense Factory Playbook for Agentic Cyber Defense
OpenAI is sharing a Defense Factory playbook that connects security tools, skills, reproducible environments, and agents into a continuous loop for discovering, validating, assigning, remediating, and independently verifying vulnerabilities.
Prompting Claude Fable 5: What to Change and What to Delete
Prompt Claude Fable 5 with less, not more: brief instructions now beat enumerated rule lists, and old skills written for prior models can degrade output. Use effort as your main cost control, expect longer turns, ground progress claims in tool results, and configure fallback to Opus 4.8 for safeguard refusals.
Prompting Claude Opus 5: Trim the Verbosity, Delete the Verification
Claude Opus 5 needs opposite prompting from its predecessors: you prompt for conciseness because effort no longer controls visible length, and you delete verification and double-check instructions because the model already does both. Constrain scope on narrow tasks, cap subagent spawning for cost, and keep thinking enabled at low effort rather than disabling it.
Prompting GPT-5.6: Message Roles, Effort, and Agentic Prompts That Work
GPT-5.6 rewards precise, explicit prompts: structure developer messages as identity, instructions, examples, then context, keep stable content first for prompt caching, and pick reasoning effort deliberately (xhigh for complex multi-step work). Move saved prompt objects into code before OpenAI shuts down v1/prompts on November 30, 2026.
Codex Record & Replay: Turn a Demonstrated Mac Workflow Into a Reusable Skill, Now in the EU, UK, and Switzerland
Record & Replay, added to the Codex app in version 26.616 on June 18, 2026, turns a workflow you demonstrate on your Mac into a reusable skill. As of September 2026 it is available in the EU, the UK, and Switzerland (since July 31), requires Computer Use to be enabled, and runs inside the ChatGPT desktop app.
Cursor Origin Enters Early Beta: Code Hosting, Pull Requests, and GitHub Sync
Cursor's Origin, a Git forge built for agents, began rolling out in early beta on all paid plans on August 17, 2026 and is still in early beta. It hosts repositories, pull requests, code browsing and search, and GitHub mirroring, with Cursor agents in every repo. Free plans are excluded, and enterprise admins can opt out.
OpenAI Codex adds Locked computer use on Mac (keep Computer Use running after lock)
Codex can keep using Mac apps after your screen locks. Locked use, now set in the ChatGPT desktop app under Settings > Computer Use, is a narrow unlock path for active Computer Use turns started from a connected device, with a short-lived authorization window and a relock on local input.
Codex Chrome Extension, Now the ChatGPT Browser Extension: How It Drives a Signed-In Browser for LinkedIn, Salesforce, Gmail, and Internal Tools
OpenAI's Codex Chrome extension is now the ChatGPT browser extension. It lets ChatGPT Work or Codex use your signed-in Chrome, Edge, Brave, Opera, or Vivaldi for sites like LinkedIn, Salesforce, and Gmail. Set it up in the ChatGPT desktop app under Settings > Computer Use, invoke it with @Chrome, and approve each new site.
Codex CLI 0.128.0 Lands Persisted `/goal` Workflows: Ralph-Style Agents That Don't Stop Until Done
OpenAI shipped Codex CLI 0.128.0 on April 30, 2026 with a persisted `/goal` system that keeps a goal alive across turns and runs the agent until it is achieved. As of September 2026 the feature is stable and on by default (`features.goals`, no config change needed), Goal mode left experimental status on May 21, 2026, and the `/goal` command is documented with set, view, edit, pause, resume, and clear subcommands. The five goal PRs in the release were authored by Eric Traut.
Perplexity Personal Computer: Complete Guide to the Always-On AI Agent for Mac and Windows
Perplexity Personal Computer is the agent built into Perplexity's desktop app: it reads and edits local files, drives native apps and the Comet browser, and runs 24/7 on a Mac mini. It needs macOS 15 or later or, since July 28, 2026, Windows 10 or 11, plus a paid Perplexity plan.
OpenAI Codex Goes 'For Almost Everything': Mac Computer Use, Browser Comment Mode, and Thread Automations Explained
OpenAI's April 16, 2026 Codex update pushed it past coding: Computer Use drives Mac apps (Windows followed on May 29), an in-app browser takes comments on rendered pages, and thread automations wake a thread on a schedule. Since July 9, 2026, Codex lives in the ChatGPT desktop app, where thread automations are scheduled tasks in a chat.
Perplexity Personal Computer vs OpenAI Codex Computer Use vs Claude Computer Use: Which AI Should Run Your Mac?
Pick Perplexity Personal Computer for an always-on desktop agent on Mac or Windows; Max includes 10,000 Computer credits a month. Pick Codex Computer Use if you live in ChatGPT and want one agent driving code, browser and desktop apps. Pick Claude Computer Use to wire it into your own harness on any OS, or in Claude Code.
Claude Code Routines: Schedule, API, and GitHub-Trigger Your AI Agents
Claude Code Routines is Anthropic's new way to run saved Claude Code configurations automatically -- by schedule, API call, or GitHub event. Routines run on Anthropic's cloud infrastructure with a prompt, repo, and MCP connectors. Available in research preview on Pro, Max, Team, and Enterprise plans.
Scheduled AI Coding Agents in 2026: Claude Code Routines vs Cursor Automations vs Codex vs Warp vs Gemini CLI
Claude Code Routines and Cursor Automations both ship schedules, API calls and native source-control events, and Cursor's trigger set is widest. ChatGPT scheduled tasks add Gmail, Slack and GitHub pull-request triggers but no API or webhook. Warp orchestrates other vendors' agents. Gemini CLI has no scheduler and lost its free individual tier on June 18, 2026.
AI Tools Landscape: What Changed in Early 2026
Three shifts defined AI tooling in early 2026: MCP settled as the cross-tool standard, coding assistants grew past autocomplete into multi-file workflow partners, and narrowly autonomous agents reached production on supervised tasks. MCP's tipping point came in June 2025, when Visual Studio Code made MCP support generally available.