AI Catchup

Claude Science Beta: An AI Workbench for Reproducible Research

By 11 min read

Claude Science is Anthropic's beta research app, not a model. It runs Python and R analyses in a local sandbox, queries 60+ scientific databases, submits jobs over SSH or to Modal, and attaches code, environment, and conversation provenance to every artifact. As of September 2026 it runs on macOS, Windows, and Linux for Pro, Max, Team, and Enterprise plans.

Anthropic launched Claude Science in public beta on June 30, 2026: a desktop research app, not a model. It runs analyses, queries scientific databases, and traces every step from data wrangling to publication, attaching the code, environment, and conversation history to each artifact it produces. As of September 2026 it runs on macOS, Windows, and Linux for Pro, Max, Team, and Enterprise plans, and it is local-first: your projects, artifacts, and raw data stay on the machine where it is installed.

The framing matters: this is a workbench wrapped around the model, not another chat surface. For how that fits Anthropic's broader push to turn Claude into a place where work happens rather than a box you paste into, see our coverage of Claude Opus 4.7 and the Claude design launch. Academic and nonprofit labs can get the app through the discounted Claude Team plan for scientists. For the team angle, our internal AI workspaces playbook covers how groups stand up shared, governed AI environments.

Key Takeaways

  • It is an app, not a model. Claude Science is a research environment that orchestrates analysis, database queries, and compute. It uses the same Claude models your plan includes.
  • macOS, Windows, and Linux. Windows 11 x64 support shipped in version 0.1.47 on September 10, 2026; the app launched on macOS and Linux only. The feature is in beta and under active development.
  • Plan gating. On for Pro and Max with no admin action. Off by default for Team and Enterprise until an Owner or Primary Owner turns it on in Organization settings. Not available on the Free plan.
  • Your infrastructure, your data. Projects, artifacts, and history live on your computer; Anthropic does not sync them to your account. Prompts and responses still go to Anthropic under its standard retention.
  • Reproducibility is the headline. Figures, tables, and notebooks carry attached code, environment, and conversation provenance, and a built-in reviewer flags claims that do not match the execution record.
  • Compute you already own. Jobs run on a workstation or HPC login node over SSH (SLURM via sbatch), or on a Modal account that bills you directly; the app connects to 60+ scientific databases.
  • Usage comes out of your normal quota. Claude Science counts against each member's standard weekly limit and uses the same seat as claude.ai.

What Claude Science Is

Claude Science is, in Anthropic's words, "your research partner for rigorous science" that "runs analyses, searches databases, and traces every step from data wrangling to publication." The distinction worth internalizing up front: this is a beta app, not a model. Anthropic's product page says it "uses the same Claude models your plan includes. What's new is everything around them: the scientific tools, database connections, and compute integrations that let Claude run full analyses on your own infrastructure."

That changes the unit of work. Instead of pasting context into a chat and copying results back out, the analysis, the data it touched, the compute it ran on, and the conversation that drove it live in one place and stay reconstructable afterward. Claude writes and runs Python, R, and shell commands (PowerShell on Windows) in a persistent kernel that keeps variables, dataframes, and loaded models in memory across the steps of a session; the kernel ends after about 30 minutes idle or when the session ends.

Availability and Plan Gating

As of September 2026, Claude Science is in beta on three platforms. Windows was added in version 0.1.47 on September 10, 2026; the June 30 launch covered macOS and Linux only.

PlatformRequirementStatus
macOS (Apple Silicon and Intel)macOS 13 or laterAvailable since launch (June 30, 2026)
WindowsWindows 11 x64; Microsoft Visual C++ Redistributable (x64)Available since version 0.1.47 (September 10, 2026)
Linuxx64 on a glibc-based distribution; bubblewrap 0.8.0 or later and socatAvailable since launch (June 30, 2026)

All three need about 5 GB of free disk space for the runtime and starter environments. Plan access, from Anthropic's admin docs:

PlanClaude Science access
Pro and MaxOn; no admin action needed
TeamOff by default; turn on in Organization settings
EnterpriseOff by default; turn on in Organization settings
FreeNot available

On Team and Enterprise plans the switch lives in Organization settings > Claude Science, and only an Owner or Primary Owner can flip it; the Admin role can assign seats but cannot turn the app on. Organizations with HIPAA compliance enabled also start with the app off, and Anthropic says usage is not covered under the Business Associate Agreement. Because it is in beta, admins should review the documentation before a broad rollout.

Active scientists at academic and nonprofit institutions can get the app through the discounted Team plan for scientists, verified through the group's principal investigator; our write-up covers eligibility and pricing.

Installing It

Anthropic's Get started doc lists a downloadable installer for macOS and Windows and a shell script for Linux; as of September 15, 2026 it does not mention Homebrew. On macOS you download the DMG for your chip and double-click it; the app then sets up its runtime and starter Python and R environments and opens in a browser tab. On Windows the signed installer adds Claude Science to the Start menu and opens the app in its own window; the first launch downloads a roughly 150 MB app window engine. On Linux you install bubblewrap and socat, run the install script from claude.ai/install-claude-science.sh, then start claude-science serve.

Two placements go beyond a laptop. The Linux command-line version can run on a cloud VM or lab server, with the app used from your own browser through an SSH tunnel; Anthropic says setup takes about five minutes. On Windows, the Linux version also runs under WSL 2 with Ubuntu 24.04 or later, though the Windows app itself does not need WSL.

Sign-in uses your Claude account; no API key is required. A setup wizard then walks through connectors, skills, which websites Claude may reach, and whether memory is on.

Reproducibility and the Reviewer

The reproducibility story is the most concrete part of the pitch. Artifacts (figures, tables, notebooks) carry full provenance: the exact code, the environment it ran in, a plain-language description of what was done, and the conversation that led there. Anthropic's pitch is that results are reproducible months later, by anyone on your team.

On top of that sits the reviewer, which Anthropic's docs describe as "a built-in verification step that independently re-reads Claude's recent responses, the approved plan, saved artifacts, and the execution record, then checks whether its claims match what ran." It runs automatically after responses and periodically during long work, and you can trigger it with Request review. Its examples include a result reported as computed when nothing ran, a value that contradicts the file it came from, a citation that does not support the claim, and a DOI that resolves to a different article. Since version 0.1.41 (August 27, 2026), findings appear as cards directly under the message they refer to.

Two limits are worth knowing. Auto-review is on by default on Max, Team, and Enterprise plans and off by default on Pro, where you turn it on per session. And the reviewer checks claims against the record; it does not re-run analyses and does not judge whether the method was the right one. Anthropic's own caution: "Verify results before relying on them in research, publication, or downstream decisions," and the app "isn't intended for clinical or diagnostic use."

Compute, Databases, and Job Submission

Claude Science is compute-aware rather than compute-blind. It manages work across a local laptop, Linux boxes, HPC login nodes, and GPU clusters, scaling "from one GPU to hundreds" and writing batch scripts as needed. Two remote paths exist:

  • SSH hosts. You add a workstation or HPC login node from your existing ~/.ssh/config; the app authenticates with your key or ssh-agent and installs nothing on the host. SLURM clusters receive jobs via sbatch, jobs survive connection loss, and each job needs approval on a card that shows the command and script. The default job timeout is 30 minutes. Remote jobs run outside the sandbox, as your user on the host.
  • Modal. Jobs run on a Modal account you own; Modal bills you directly and Anthropic never sees a payment method. Each job is approved with its machine spec and maximum billable time visible. There is no spend ceiling in the app; the default container timeout is 12 hours (maximum 23), and closing the app does not cancel a running job.

It also connects to 60+ scientific databases so you can query them without learning each interface, and ships pre-configured for genomics, single-cell analysis, proteomics, structural biology, and cheminformatics. Native viewing covers proteins, structures, molecules, alignments, genomic tracks, chemical structures, and PDFs. Any pipeline can be saved as a reusable skill, any Model Context Protocol server can be added as a custom connector, and NVIDIA BioNeMo NIM model endpoints (Evo 2, Boltz-2, OpenFold3) connect from Settings > Compute on macOS and Linux, not Windows.

What Stays Local and What Anthropic Receives

Anthropic's data doc calls the app "local-first": conversation history and artifacts are stored on the member's computer and are not synced to the Claude account or other devices, so a second computer starts fresh. What does travel is each prompt and Claude's response, logged under Anthropic's standard retention policy for model traffic, plus product telemetry and redacted error reports that device configuration can turn off. Code and data sent to your own SSH hosts or Modal account go directly there and do not pass through Anthropic.

For Enterprise organizations with the Compliance API enabled, Anthropic also retains session transcripts (in beta) that the compliance team can retrieve; they include prompts, responses, tool calls, and the text of files Claude read, but not extended thinking or the system prompt. Admins can turn SSH hosts, Modal, scientific model endpoints, custom connectors, and memory off for the organization.

When to Reach for Claude Science

You want to...Claude Science fit
Run reproducible analyses with attached provenanceStrong; artifacts carry code, environment, description, and history
Keep raw datasets and compute on your own infrastructureStrong; local-first, with SSH and Modal for heavier jobs
Submit and manage HPC or GPU jobs from one placeStrong; SLURM via sbatch, workstations, or Modal, each job approved
Query many scientific databases without bespoke toolingStrong; 60+ databases plus MCP custom connectors
Work on a Windows PCAvailable since September 10, 2026; PowerShell, local NTFS or ReFS folders only, no BioNeMo NIM endpoints
Roll out org-wide without admin reviewWeak; beta, off by default on Team and Enterprise, Owner-level switch
Use it for clinical or diagnostic decisionsNot intended; Anthropic says so in the docs

Why This Matters

The interesting move here is positioning. Most AI-for-science effort goes into better models; Claude Science instead invests in the environment around the model: provenance, compute management, database connectivity, and a reviewer that catches claims the execution record does not support. That is the part of the research workflow that chat surfaces leave to the user, and it is where reproducibility actually breaks down.

The honest caveat is the one Anthropic states: this is beta, it may change, the reviewer reduces but does not eliminate errors, and on Team and Enterprise it is gated behind an Owner-level switch. Treat it as an early but unusually concrete attempt to make AI-assisted research traceable end to end, not a finished platform.

FAQ

See the structured FAQ in the schema header for question-level details: what Claude Science is, which platforms and plans support it, whether it runs on Windows, where data and compute run, how it supports reproducibility, and how usage is billed.

Sources

Related coverage

Frequently Asked Questions

What is Claude Science?

Claude Science is a desktop research app from Anthropic, in beta, that pairs Claude with an analysis environment on your computer. You describe a task in plain language; Claude writes and runs Python, R, or shell code in a sandbox, pulls data from scientific databases through connectors, and saves results as versioned artifacts with a provenance record.

Which platforms and plans support Claude Science?

As of September 2026, Claude Science runs on macOS 13 or later (Apple Silicon and Intel), Windows 11 x64, and x64 Linux. Windows support arrived in version 0.1.47 on September 10, 2026. It is available on Pro, Max, Team, and Enterprise plans; Team and Enterprise organizations must turn it on first, and the Free plan cannot use it.

Does Claude Science work on Windows?

Yes, since version 0.1.47 on September 10, 2026. The Windows app needs Windows 11 x64 and opens in its own window rather than a browser tab. Shell commands run in PowerShell, only folders on local NTFS or ReFS drives can be granted, and the NVIDIA BioNeMo NIM model endpoints are not available on Windows.

Where does my data and compute run?

Claude Science is local-first: projects, artifacts, and conversation history stay on the computer where it is installed, and Anthropic does not sync them to your account. Code runs in a local sandbox. Jobs can run on an HPC login node over SSH or on a Modal account you pay for. Prompts and responses still go to Anthropic.

How does Claude Science support reproducibility?

Every artifact carries the exact code, environment, plain-language description, and conversation that produced it. A built-in reviewer re-reads Claude's responses, the approved plan, saved artifacts, and the execution record, and flags claims that do not match what ran. Auto-review is on by default on Max, Team, and Enterprise plans and off by default on Pro.

Does Claude Science count against my Claude usage limits?

Yes. Anthropic's docs say Claude Science usage counts against each member's standard weekly quota and uses the same seat as the rest of claude.ai. On Pro and Max plans, sessions pause and ask before spending extra usage credits, and Settings > Usage shows the credit balance and monthly spend limit.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.