Benchmarks Articles
6 articles across AI Catchup's news, guides, tutorials, and comparisons.
All Benchmarks articles
Anthropic Releases Claude Haiku 5.5, Its Fastest Model and First Haiku With Effort Levels
Anthropic released Claude Haiku 5.5 on October 7, 2026, its fastest model to date and the first Haiku with effort levels. It scores 72.4% on OSWorld 2.1 against Haiku 4.5's 15.7%, and costs $0.10 per million input and $0.50 per million output tokens for prompts up to 100,000 tokens, 90% less than Haiku 4.5.
GPT-6.1 Sol: Astra-Level Coding at a Fifth of the Price, With Cost at Every Effort Level
OpenAI released GPT-6.1 Sol on September 29, 2026, at $2 input and $10 output per million tokens, a fifth of GPT-6 Astra's price, with cached input at $0.10. On OpenAI's charts it edges past Astra's best DeepSWE score at high effort and trails Astra by 2.1 points on OSWorld offline. The API model ID is gpt-6.1-sol.
Claude Opus 5.5: Fable 5.1-Level Scores at $4 and $20, With Benchmarks, Effort Costs, and Migration Guide
Claude Opus 5.5, released September 22, 2026, beats Claude Fable 5.1 on every benchmark Anthropic published and costs $4 and $20 per million tokens, 60% less than Fable 5.1. It scores 66.4% on Terminal-Bench 4.0 against 57.9% for GPT-6 Astra. Thinking is always on, the default effort is medium, and four API changes can break Opus 5 code.
GPT-6 Sol and Luna: Half the Price of GPT-5.6, With Benchmarks and Cost at Every Effort Level
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, at $2 and $10 and $0.10 and $0.50 per million tokens, half the price of GPT-5.6. Luna scores 66.6% on DeepSWE for about $0.22 a task. Since October 7 the pair has been rolling out as ChatGPT's Chat models, and GPT-6.1 Sol has succeeded Sol at the same price.
Google Ships Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google released three Gemini models on July 21, 2026: Gemini 3.6 Flash, a coding and knowledge-work workhorse at $1.50/$7.50 per million tokens that uses 17% fewer output tokens; Gemini 3.5 Flash-Lite, a 350 tokens-per-second model at $0.30/$2.50 for high-throughput agents; and Gemini 3.5 Flash Cyber, a security-specialized model piloted with governments and trusted partners.
OpenAI introduces GeneBench-Pro, a research-level benchmark for agentic computational biology
On June 30, 2026, OpenAI announced GeneBench-Pro: a research-level benchmark meant to measure how well AI agents navigate messy biological data and make the judgment calls real computational biology depends on. OpenAI says GeneBench-Pro contains 129 questions and is open-sourcing 10 representative case studies as a public package on Hugging Face under the MIT License.