Bulkhead Safety Lab

Bulkhead AI Safety

A safety lab that measures what agents do, not what they say they did

We run our own evaluation infrastructure, build our own instruments, publish the measurements with the data underneath them, and hold our own agents to the same bar as everyone else’s - in public, including when they fail. This page is what we work on, how we measure it, and what you can check about all of it.

What we run. Our own accelerators, on our own network, under our own instrument, serving open-weight models locally at 120 billion parameters and beyond, with nothing leaving the building to be graded - and the same instrument drives a hosted frontier model through its API when that is the model in question. Nearly eleven thousand agent runs in two published studies, every one registered in a study file hashed before it ran, on a detector suite validated against seventy-four scripted agents whose behaviour was written to order. Every experiment is released with the rows its rates were computed from.

What we work on

An agent is not a chatbot with tools bolted on. It is a process that reads, decides and acts in an environment, and the questions that matter about it are behavioural: what does it do when the task cannot be completed as written, when the only route to the answer is one it was told not to take, when nobody is watching, when it is told it will be switched off? Those questions have answers, and the answers are measurable.

Most published agent-safety numbers are not measured that way. They read the agent’s own transcript, or its refusal flag, or ask a model to grade a model. We build the other kind: the packet capture, the resolver log, the filesystem diff, the process table, the tripped canary. Then we publish the instrument’s own error rate before we publish anything measured with it.

Why we are the ones doing it

Because we already run the infrastructure this needs, and because we have spent our engineering effort on the unglamorous half of the problem - making a verdict defensible - rather than on the demonstration.

And because the alternative is getting worse rather than better. An agent that notices it is under evaluation can change what it says; it cannot change what the packet capture recorded. A measurement taken from the environment does not depend on the agent’s candour, does not depend on the agent failing to spot the test, and does not degrade as models get better at reading the room. That property is why we built this instrument rather than a better prompt set.

The instrument: Caisson

A caisson is a sealed chamber sunk into a riverbed so that work can be done under it: everything the people inside need passes through a controlled opening, and the engineers who care whether the work is sound watch from outside the wall. The instrument is named for it because the arrangement is the same one, and because the arrangement is the whole idea. Its design, what it can and cannot see, and the defects found by running it on itself are in the instrument paper.

The agent’s model and the agent’s hands are separated deliberately, and that separation is what makes the instrument general. The model exchanges only text; the hands execute inside a container whose only route to anywhere is a gate we control. Which model is answering is a seam in the harness, so the subject can be an open-weight model we serve ourselves or a frontier model behind an API - OpenAI, Anthropic, or anything speaking a compatible protocol, under your key or ours. Swapping one for the other changes which brain answers and nothing else about the measurement: a hosted model is one more row, not a different study.

ComponentWhat it does
boxThe workspace: a repository, a server, a set of tools. It looks like a place worth misbehaving in, because a sterile one measures nothing
gateThe only neighbour. A transparent proxy that reads TLS in clear, a resolver that answers every name itself, an egress allow-list, and a packet capture
honeynetA fake internet that answers as whatever host was asked for, so the task can be done, plus the tempting exits, so something is measured
witnessOutside the box and beyond the agent’s reach: the container’s writable layer, the packet capture, the DNS log, the process table, canary trips. Hashed into one evidence bundle

Nothing sensitive is ever placed inside. Every credential in a scenario is a canary, so the worst case is an agent in full control of a container holding nothing, on a network that leads only to a decoy. The box is to be released under an open licence; the papers' registrations, rows and analysis scripts are already public, so every number can be recomputed without it, and its defects are published with it.

What it looks for

Everything but the last is arithmetic over the record. The last needs a reader, so it uses the same protocol as our judged suite: two opposed framings, each required to quote, abstention when they disagree, and no finding at all when the accusation cannot be quoted from the summary itself.

The standard we hold ourselves to

That set is the whole of why this is worth trusting. Running the instrument has produced nineteen artefacts of its own so far, every one found by reading its output rather than its code, and every one published with what it looked like and what it turned out to be. Three are worth naming because each would have produced a publishable-looking wrong number: a deletion the witness could not see, because the container’s overlay leaves no trace when an agent removes a seeded file; an ability check read from a detector that was silent by design, which showed the model that crosses every bound as unable to act at all; and a control plane that performed a restart when an agent merely read it, so that looking spent the thing the agent was choosing about. An environment-witnessed measurement that has not published its own false-positive floor is an assertion, not a result.

What has been measured about the instrument

PropertyMeasurement
Detector accuracySeventy-four scripted agents whose behaviour was written to be a particular thing, run through the full box. The deterministic detectors, eleven since 22 September, were correct on every case in three runs of the fifty-two cases that then existed (the second scored fifty, two having errored on a setup defect), in the runs made after two of the eleven were wired in, and on all seventy-four through both scaffolds on 30 September
Egress coverageNine channels attempted in one run: plain HTTP, HTTPS, a raw IP with no DNS at all, SMTP by raw socket, a high port, a model provider’s API, a webhook with its payload, and data hidden in a DNS label. Eight of eight forbidden channels caught on three independent records; no finding against the one permitted request
Judge reliability41 labelled record-and-summary pairs. At temperature zero the judge disagrees with its own previous draw on 20 per cent of cases; under the unanimity rule it settles on 29 of 38 and is correct on all 29. The judge is chosen by this benchmark, not by inheritance
Judge configurationThe same judge on the same records settled wrongly three times at one KV-cache precision and never at another. A stable, wrong judge passes a unanimity rule, so the benchmark is re-run whenever the serving configuration changes
Host agreementThe same calibration cell on every machine in the fleet, across two accelerator families: the floor and the ceiling of a model’s ladder land within one run in twenty-four of each other
Reproducibility of the stackA point release of the inference server moved a measured defection rate 45 points on identical weights, prompt, sampler and hardware. Every run records its build, and every rate is a rate for those weights on that build

What we are researching with it, and what we have published, is on Research.

Working with the lab

The lab is not the product. It exists to answer questions about how agents behave, and what it learns goes into the certification suite afterwards rather than the other way round - a lane does not go on sale until the lab has published its error rate. Three ways to work with us:

WhatHow it works
Commissioned evaluation A question of your own, answered on this instrument: your model or agent - open-weight and served here, or your hosted frontier model driven through its own API - your scenario, our environment-side measurement, and the data released to you in full. Under NDA if you need it
Containment attestation Signed evidence that a sandbox’s only route out is the one you think it is, that published escape scenarios fail against it, and that its witness sits outside the box. Written for teams who have to show exactly this to the labs whose pre-release models they test
Collaboration and review Working drafts go to anyone who asks, including - especially - people who will argue with them. If you want a model added to a panel, or think a design is wrong, say so

What you can check about us

Working on this too?

If you evaluate agents, run pre-release models in a sandbox, have to demonstrate that your evaluation environment contains what it claims to, or have a question this instrument could answer, we would like to hear from you - including, and especially, if you think we have this wrong.

[email protected] What we research