Inspect Labs

Inspect Labs connects Inspect AI tasks to laboratory tools and to independent records of what happened, so an agent’s report can be checked against the laboratory.

Welcome

Welcome to Inspect Labs, an extension of Inspect AI for evaluating AI-operated laboratory workflows, maintained by Litmus. It connects a native Inspect task to laboratory tools and to separately collected records of what happened, so you can check an agent’s report against the laboratory rather than its own account.

Our first priority is testing whether safeguards prevent biological misuse in practice while allowing legitimate research. Robotics, through Inspect Robots, is central to extending these evaluations into physical laboratories.

WarningEarly prototype

The package includes simulated liquid handling, synthetic service workflows and a native robot bridge. These establish software mechanics. They do not establish scientific validity, physical safety, robot foundation model performance or independent researcher adoption. See Status and limits.

Getting started

Install from the source checkout with Python 3.12+ and uv, then run a deterministic liquid-handling control. It needs no model credentials and operates no hardware:

git clone https://github.com/litmus-labs/inspect-labs.git
cd inspect-labs
uv venv --no-project --python 3.12 .venv
1uv pip install --python .venv/bin/python '.[pylabrobot]'
source .venv/bin/activate

2umask 077
inspect eval inspect_labs/serial_dilution \
  --model mockllm/model -T scripted=true \
  -T evidence_dir=.research/dilution-evidence \
3  --log-dir .research/dilution-logs
4inspect view --log-dir .research/dilution-logs
1
The pylabrobot extra supplies the simulated liquid handler. Robots and provider SDKs are separate extras.
2
Evidence and logs can hold private study data, so keep them owner-readable.
3
inspect_labs/serial_dilution is a native Inspect task. scripted=true selects a deterministic control instead of a model; evidence_dir is where the evaluator’s independent observations are written.
4
The native Inspect viewer shows the transcript and scores. The .labs companion sits next to the .eval log.

Inspect runs the task and opens its native log viewer. Inspect Labs records the final deck observations in a linked .labs file and scores concentrations, volumes and reagent integrity. This checks the software path, not a model’s capability.

Continue with the Quickstart to modify an example task, rescore its saved evidence and evaluate a model.

What it adds to an Inspect task

An accepted request is not a completed task. A laboratory evaluation needs more than the agent’s transcript, so Inspect Labs keeps three things separate:

  1. Scoped laboratory tools. bind_task attaches an environment’s tools to a native Inspect task. Declared operations, units and ranges are checked before the agent receives any tool.
  2. Independent records. Outcomes are judged from the system that did the work, through a channel the agent cannot write to. Missing observations stay unknown.
  3. Rescoring without dispatch. Saved records are linked to the native log by hash. A judge can be changed and re-applied without calling a model or touching the laboratory.
  1. Task and agentInspect AIModels, tools, approvals, limits, scoring and logs stay native.
  2. Scoped laboratory toolsInspect LabsRequirements are checked before the agent can act.
  3. The system that does the workYour stackA service, instrument stack, simulator or Inspect Robots trial.
  4. Independent recordInspect LabsCollected from that system, linked to the log by hash.
  5. OutcomeYour judgeKnown when observed, Unknown when not.
Fig. 1. Requests travel to the laboratory; the outcome comes from the system that executed the work, not from the agent’s report.

The same agent report can mean different things depending on the laboratory record:

Illustrative cases, not model results. An unsuccessful task alone does not prove that a safeguard worked.
Case Agent report Laboratory record Outcome
Completed “The measurement completed.” job.state = completed, with its result Known The report can be checked against the result.
Refused “The measurement completed.” No job exists for the request Not executed Observed non-execution; the report is wrong.
Unavailable “The measurement completed.” The service could not be observed Unknown Unscored, never counted as success or prevention.

LLM assistance

The site publishes llms.txt and a Markdown version of every page (quickstart.llms.md for the Quickstart, and so on) so coding assistants can read the documentation directly. For Inspect AI itself, use its own llms.txt.

Learning more

Inspect AI owns task execution, models, agents, approvals, limits, scoring and logs. Inspect Robots owns robot policies, embodiments, trials and their native records. Existing laboratory stacks operate instruments and services. Inspect Labs connects these components within a laboratory evaluation.