The AI Catchup - July 22, 2026
Welcome back. Last issue was about surfaces: ChatGPT split into three agents and Cursor shipped branching threads. This issue is about headroom. Cursor doubled the usage limits on every individual and teams plan, Claude Code gave iOS developers the live app-testing loop they have been asking for, and Google priced its two new Gemini models to make high-volume agent work cheaper. Every lead story this week either raises your ceiling or lowers your bill.
Let us get into it.
At a Glance
| What shipped | Why you care | Can you use it today? |
|---|---|---|
| Cursor doubled usage limits | 2x usage on all individual and teams plans, covering Grok, Composer, and any new Cursor models | Yes. Live now, no end date announced |
| Claude Code iOS Simulator pane | Watch Claude build, run, and tap through your iOS app live, next to the conversation | Public beta on macOS for Pro, Max, and Team plans |
| Gemini 3.6 Flash + 3.5 Flash-Lite | A cheaper, more token-efficient coding workhorse and a 350 tok/s agent model | Yes. API, AI Studio, and the Gemini app since July 21 |
| Codex CLI 0.145.0 | /import migrates Cursor and Claude Code setups; 272k context restored for GPT-5.6 | Yes. Update the CLI (July 21 release) |
| GPT-Red | OpenAI's internal automated red-teamer, and the reason GPT-5.6 is harder to prompt-inject | No. Internal-only by design |
This Week in AI
Cursor made the most concrete move of the week: usage limits are now doubled on all individual and teams plans (announced July 21), covering Grok, Composer, and any new Cursor models. That last clause is worth reading twice, because it means future model additions inherit the doubled allocation automatically. What began as a one-week Grok 4.5 launch promo is now the plan allocation, with no end date announced.
The timing writes its own headline. Anthropic's +50% Claude Code weekly-limits promo expired July 19, and two days later Cursor doubled, permanently as far as anyone can tell. If you budget agent work across tools, this week's scoreboard is simple: Cursor plans carry double the first-party usage, and Claude Code limits are back to standard. Practically, that means two things. If you throttled Composer or Grok usage to stay inside your allocation, the ceiling you tuned for is twice as high now. And if you routed overflow work to a bring-your-own-key model to conserve pool usage, re-run that math, because the pool may be the cheaper path again.
Cursor also pushed its agent deeper into Slack this week (changelog, July 17). It now shares a plan before it starts so you can redirect it early, targets named multi-repo environments instead of assuming one default repo, and can read from and post to other Slack channels and threads, pulling context from where the discussion happened and reporting back where the stakeholders are.
Claude Code Can Now Watch Your iPhone App Run
The feature of the week for anyone shipping mobile: Claude Code Desktop added an iOS Simulator pane in public beta. When Claude builds, runs, or checks your iOS app, Apple's Simulator opens in a pane next to the conversation. Claude installs the app, taps through it, and reads the screen to verify its own changes while you watch. The pane is interactive, so you can grab the device yourself mid-session, navigate to a screen, and ask "does this look right?"
The design details are what make it usable. It drives the simulator directly, with no computer use, no screen takeover, and no macOS Accessibility permissions, so you can keep working while Claude tests. Consent is per device with a visible "Claude is using this device" badge, each session gets its own devices (up to four panes), and Claude shuts down simulators it booted when they are no longer in use. Requirements: macOS, Claude Desktop v1.24012.0+, Xcode, on Pro, Max, and Team plans. Until now, the agent could edit Swift all day but checking the UI meant screen-control hacks or you playing screenshot courier. The build, launch, tap, read, fix loop now runs in one pane.
Google Dropped Two Gemini Models Priced for Agent Work
Google shipped a triple model release on July 21, and the two you can actually use are priced to move. Gemini 3.6 Flash is the new coding and knowledge-work workhorse: 49% on DeepSWE (vs 37% for 3.5 Flash) at $1.50/$7.50 per million tokens. The number to anchor on is that it uses 17% fewer output tokens on average, and up to 65% fewer on DeepSWE. Agentic bills scale with output tokens, so it can be dramatically cheaper per completed task before you even compare rates. Gemini 3.5 Flash-Lite streams at 350 tokens per second for $0.30/$2.50 and posts 54% on Terminal-Bench 2.1, a pointed play for the tool-calling and orchestration tier where fast, cheap models like GPT-5.6 Luna compete.
The third model, Gemini 3.5 Flash Cyber, is not for sale: a vulnerability-hunting model piloted exclusively with governments and trusted partners inside Google's CodeMender agent. It fits a pattern this week. See Quick Hits for OpenAI's version.
Ship It This Week
Three direct wins you can act on today.
Un-throttle your Cursor usage. If you rationed Composer or Grok to stay inside your allocation, the ceiling doubled. Remove the self-imposed limits and re-test your real workflow at full speed.
If you build iOS apps, put the simulator pane on your most fragile flow. Update Claude Desktop, open the app project, and ask Claude to run it in the simulator and tap through signup. Watch the first run, since the device badge shows exactly when Claude is driving, and keep test accounts only on devices Claude uses.
Run one workload through Gemini 3.6 Flash and measure output tokens. The 17% efficiency gain only shows in your own traces. Compare cost per finished task against your current model, not sticker price.
Quick Hits
- Codex CLI 0.145.0 will import your existing setup. The July 21 release (changelog) expands
/importto migrate Cursor and Claude Code settings, MCP servers, plugins, sessions, commands, and project-scoped memories into Codex, so trying Codex no longer means rebuilding your tooling from scratch. The release also adds searchable thread history with sub-agent support, stabilized multi-agent V2, and Amazon Bedrock login. Worth updating even if you skip all of that: the July 18 patch restored the full 272,000-token context windows for GPT-5.6 Sol, Terra, and Luna. - OpenAI showed off GPT-Red, and the numbers are startling. Its internal automated red-teaming model succeeded in 84% of scenarios versus 13% for human red-teamers on an indirect prompt injection benchmark, and OpenAI credits the attack-into-training loop for GPT-5.6 Sol's 6x fewer injection failures. Like Google's Flash Cyber, it is deliberately kept off the public shelf. The attack surface it targets (emails, webpages, tool responses, code repositories) is what your coding agent reads all day, so keep your own layer configured with least-privilege setups like Codex permission profiles.
- Claude Code shipped a screen reader mode. An opt-in mode (details) replaces boxes, spinners, and redraws with labeled linear text that VoiceOver and NVDA read in order: searchable
you:/claude:/Permission Required:labels, numbered menus, and a terminal bell when Claude needs you. It is the first major coding agent with a documented, versioned screen reader mode. - NotebookLM is now Gemini Notebook. Google renamed its research tool (July 16), citing 30 million+ users and 600,000+ organizations. The substantive part: the Gemini 3.5 + Antigravity upgrade comes to AI Pro subscribers over the coming weeks, and notebooks are headed for Search's AI Mode.
- ChatGPT Voice now runs on GPT-Live-1. The new Voice listens while it speaks, uses web search and memory mid-conversation, and handles text and images in one thread. It is rolling out across consumer plans including Free. Give it a multi-step errand, not a single question; that is where the upgrade shows.
That is it for this issue. The through-line this week is capacity: double the Cursor usage, a live simulator loop that removes the slowest step in mobile agent work, and Gemini models priced so the cheap tier can carry more of your pipeline. The vendors spent the week competing on how much you can get done, not just how smart the model is, which is the competition you actually want. If a teammate is still rationing their Cursor usage against last month's limits, forward this.
Until next week, stay caught up.