AI Catchup

Search AI Catchup

Find articles, comparisons, tutorials, editor's picks, and newsletter editions across the complete archive.

Latest across AI Catchup

24 results

News & Analysis

Claude Science Beta: An AI Workbench for Reproducible Research

Claude Science is Anthropic's beta research app, not a model. It runs Python and R analyses in a local sandbox, queries 60+ scientific databases, submits jobs over SSH or to Modal, and attaches code, environment, and conversation provenance to every artifact. As of September 2026 it runs on macOS, Windows, and Linux for Pro, Max, Team, and Enterprise plans.

News & Analysis

Ox Alpha Was Z.ai's GLM-5.3-Flash, and the Free Week Is Over

Ox Alpha was Z.ai's GLM-5.3-Flash. OpenRouter's model page now names the developer, the stealth listing serves no providers, and OpenCode has dropped Ox Alpha Free. The named model lists at $0.15 per 1M input tokens and $0.50 output; the 50% launch discount ended September 9, 2026, and as of September 10 Z.ai's pricing page carries list prices only.

News & Analysis

Claude Managed Agents Adds a Terminal Session Viewer and Auto Permission Mode

Claude Managed Agents now includes a terminal workflow for attaching to live sessions and a local web session viewer. Anthropic also added an auto permission mode that evaluates each agent or MCP tool call against the intent in user.message and runs, denies, or pauses for approval.

Practices & Workflows

Master Claude Code's 1M Context Window: Rewind, Compact, Clear, and Subagents

Claude Code's 1M token context window opens longer autonomous sessions but introduces 'context rot' -- degraded performance as the window fills. Master four turn-end tools: /rewind to drop bad branches, /compact to summarize and continue, /clear to start fresh with a distilled brief, and subagents to wall off noisy work in their own context.

News & Analysis

OpenAI Launches ChatGPT for Financial Services

OpenAI introduced ChatGPT for Financial Services, a tailored ChatGPT Work experience for eligible financial institutions. It combines hosted datasets from providers such as Daloopa, PitchBook, LSEG News, and Crunchbase with GPT-6 Astra, granular citations, spreadsheet and presentation generation, firm templates, and enterprise access controls.

News & Analysis

ChatGPT Sites Adds Teammate Editing, External Viewers, and Public Publishing Controls

OpenAI’s latest ChatGPT Sites update expands the sharing model around teammate editing, named external viewers, and public publishing. A Site owner can give an active member of the same ChatGPT workspace Can edit access; editors can update and save the Site, then publish later versions to the same URL after the owner’s first publish. Owners retain access, URL, secrets, custom-domain, and audience controls.

Practices & Workflows

Prompting Claude Opus 5: Trim the Verbosity, Delete the Verification

Claude Opus 5 needs opposite prompting from its predecessors: you prompt for conciseness because effort no longer controls visible length, and you delete verification and double-check instructions because the model already does both. Constrain scope on narrow tasks, cap subagent spawning for cost, and keep thinking enabled at low effort rather than disabling it.

News & Analysis

Claude Code Adds Plugin Evals for Regression Testing and No-Plugin Baselines

Claude Code now includes `claude plugin eval`, a testing workflow for plugin and skill authors. The command can create cases and graders, run a plugin in isolated non-interactive sessions, compare results with a no-plugin baseline, generate an HTML report, and support CI gates for plugin changes. Evals require Claude Code v2.1.269 or later and consume model usage.

News & Analysis

GPT-Rosalind Leaves Research Preview for Eligible Organizations Worldwide

GPT-Rosalind, OpenAI's purpose-built life sciences reasoning model, is out of research preview for eligible organizations worldwide through OpenAI's trusted-access program. OpenAI says qualified customers can use it in the API, Codex, and ChatGPT Enterprise, while a Life Sciences research plugin connects Codex to more than 50 public tools and data sources. Published pricing takes effect October 5, 2026.

News & Analysis

Claude Platform's 'ant' CLI Brings the Full Claude API to Your Terminal

Anthropic's 'ant' CLI exposes every Claude API resource as a terminal subcommand. As of September 2026 you install it through Homebrew, curl, or Go, authenticate by browser login, API key, or Workload Identity Federation, manage agents as code with 'ant apply' (CLI 1.30.0 or later), and run a self-hosted Managed Agents worker with 'ant beta:worker poll' or 'ant beta:worker run'.

News & Analysis

Cursor Projects Give a Coordinator Agent a Long-Lived Body of Work

Cursor's Projects beta moves beyond one chat at a time: a coordinator agent keeps context over months, delegates implementation and testing to subagents, shares research and artifacts across cloud and local machines, and can watch PRs, Slack, or a schedule for recurring work.

News & Analysis

Claude Fable 5.1 and Mythos 5.1: Same Price, Cheaper Cache Reads, Fewer Safeguard Interventions

Anthropic released Claude Fable 5.1, available today as claude-fable-5-1 at $10 per million input and $50 per million output tokens, with cache reads cut 75% to $0.25 per million. Anthropic estimates typical workloads cost about 25% less than Fable 5. Claude Mythos 5.1 is the same model with fewer cyber and life-science safeguards, limited to vetted US organizations.

News & Analysis

OpenAI Shares a Defense Factory Playbook for Agentic Cyber Defense

OpenAI is sharing a Defense Factory playbook that connects security tools, skills, reproducible environments, and agents into a continuous loop for discovering, validating, assigning, remediating, and independently verifying vulnerabilities.

News & Analysis

Claude Code Desktop Adds Auto-Continue After Usage Limits

Claude Code waits in the open session and continues your task automatically when a claude.ai usage limit resets. It is on by default from version 2.1.234 in interactive sessions signed in with a claude.ai subscription. Turn it off in /config under 'Continue automatically at usage limit' or set autoContinueAtUsageLimit to false. It resumes work; it does not raise the limit.

News & Analysis

Claude Marketplace Adds Cursor, CrowdStrike, Factory, Gamma, and Vercel

Anthropic is expanding Claude Marketplace with Cursor, CrowdStrike, Factory, Gamma, and Vercel. Enterprise customers can use an existing Anthropic spend commitment to buy Claude-powered solutions from marketplace partners, with Anthropic directing buyers to their account team rather than a self-serve checkout.

News & Analysis

OpenAI Launches ChatGPT Images 2.5 and GPT-Image-2.5 API Models

OpenAI launched ChatGPT Images 2.5 with sharper image fidelity, faster generation, more precise multi-turn edits, Sketch, templates, and comment-based editing. The same release adds GPT-Image-2.5 Flare for fast everyday API generation and Sunburst for more precise, premium creative workflows.

News & Analysis

Codex Build iOS Apps Plugin: Mirror the Simulator in the Browser and Hot-Reload SwiftUI Previews

OpenAI's Build iOS Apps plugin for Codex bundles nine iOS and Swift skills. The June 2026 addition mirrors the iOS Simulator into the Codex in-app browser and hot-reloads Swift Package-backed SwiftUI previews. XcodeBuildMCP handles simulator build, run, and debug. It installs from the /plugins browser in Codex CLI or the Plugins tab in the ChatGPT desktop app.

News & Analysis

OpenAI Resets ChatGPT Work And Codex Limits For Paid Users

OpenAI executive Tibo announced usage resets for paid ChatGPT Work and Codex users on August 8 and August 29, 2026, then posted that a global reset would be made for all paid subscriptions on September 7. The approved posts do not publish a numeric allowance, a new API billing rule, or a new recurring reset schedule.

News & Analysis

OpenAI Says GPT-6 Astra Now Uses Less ChatGPT Subscription Usage on Long Tasks

OpenAI executive Tibo says recent improvements to GPT-6 Astra reduce the subscription usage drawn by long-tail work for power users signed in with a ChatGPT account. The approved announcement claims up to 3–4x less usage and says there is no change in quality, but it does not publish implementation details, a new quota, or an API billing change.