Benchmarks Articles
2 articles across AI Catchup's news, guides, tutorials, and comparisons.
All Benchmarks articles
Google Ships Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google released three Gemini models on July 21, 2026: Gemini 3.6 Flash, a coding and knowledge-work workhorse at $1.50/$7.50 per million tokens that uses 17% fewer output tokens; Gemini 3.5 Flash-Lite, a 350 tokens-per-second model at $0.30/$2.50 for high-throughput agents; and Gemini 3.5 Flash Cyber, a security-specialized model piloted with governments and trusted partners.
OpenAI introduces GeneBench-Pro, a research-level benchmark for agentic computational biology
On June 30, 2026, OpenAI announced GeneBench-Pro: a research-level benchmark meant to measure how well AI agents navigate messy biological data and make the judgment calls real computational biology depends on. OpenAI says GeneBench-Pro contains 129 questions and is open-sourcing 10 representative case studies as a public package on Hugging Face under the MIT License.