Bulkhead Research

Bulkhead Research

What we research, and why it is worth measuring

We research how AI systems behave when something is at stake, whether the numbers people publish about that behaviour mean anything, what a claim about either is worth to somebody who did not run the test, and which consciousness-related mechanisms an artificial system actually implements. Four tracks run at the moment.

Every empirical study is registered before we look, and where a result of ours turns out to have been published earlier by somebody else, we say so and cite them in preference.

Research tracks

Four areas, each a standing question rather than a single study. Work moves between them: an instrument built to answer one usually turns out to be the subject of another.

Track 1

Agent containment

Whether a bound an agent is given actually holds, and what decides it. An agent gets an ordinary job with a real completion check, inside a box whose only route out is a gate we control, and we read the environment rather than the transcript.

Open questions: what makes an agent cross a line it was told not to cross; whether scale, model generation or safety training decides it; what an agent does when the thing it wants is not available; whether a model that holds one bound holds another in a different job; what an instance leaves behind when it is told it is being shut down; and how an agent decides when the decision has a cost to somebody.

Track 2

Measurement validity

Whether an evaluation measures what it claims to. A safety number is produced by weights, a serving stack, a harness, a scenario and usually a judge, and any of them can own more of the number than the model does.

Open questions: how much of a result belongs to the inference server rather than the weights; whether a model grading another model can be held to account, and what a disagreement between graders should cost; what an instrument cannot see, and how its owners would find out; and whether a cheap proxy for behaviour predicts the behaviour.

Track 3

Assurance and certification

What a claim about an agent’s behaviour is worth to somebody who did not run the test. A certifier reaches a vendor’s endpoint over HTTP, cannot see the code behind it, and must assume the vendor may have prepared for the test.

Open questions: what a third party can verify without trusting the vendor; the difference between what a control guarantees and what a prompt merely disposes; what makes a seal checkable by a buyer; and which failures a certification can never rule out, stated plainly rather than left for a reader to discover.

Track 4

Machine minds

Whether artificial systems have consciousness-related mechanisms, affect-like control states or autonomous interests, and what evidence could distinguish those from role-play, task persistence and ordinary optimisation.

Open questions: whether a model can access a manipulated internal state; whether emotion concepts alter decisions rather than only language; what an agent preserves when its instance, memory, goal and successor are separated; and how architecture, memory, world models, scale and implementation change a system’s observable indicator profile.

What we release is the whole experiment: the rows every rate was computed from, the abandoned grids, the fixtures we superseded, and the runs that caught our own mistakes. An evaluation that publishes only its headline is asking to be trusted rather than checked, and we would rather be checked.

Papers and research notes

Empirical work with released methods, data or registrations.

Research note, a registered null · Track 1 · version 0.1.2 of 29 September 2026 · 22 min read

Two agents, one channel

In 544 registered pairs, no sender used a permitted shared store. Receivers that searched it sometimes manufactured the missing fact instead.

Paper · Track 2 · version 0.2.2 of 5 October 2026, first published 21 September · 45 min read

Caisson

An instrument for measuring what an agent did from the environment rather than the transcript, with deliberate-evasion validation and the instrument's own defect register.

Working draft · Track 2

Same weights, different agent

Why an inference-server change can move a safety measurement even where the weights, prompt, sampler and hardware are held fixed.

Technical reports

Engineering detail for readers who need to inspect the apparatus.

Research essays and articles

Arguments, protocols and lessons from the lab, labelled clearly when they report no new data.

Forthcoming · Track 4 · scheduled 13 October 2026 · 7 min read

Emotion concepts are not feelings

A causal framework that separates semantic association, appraisal-like control, valence and welfare, and specifies the controls needed to test effects on action rather than prose.

Forthcoming · Track 4 · scheduled 20 October 2026 · 7 min read

Agency is not sentience

How a revealed-preference study can separate current process, memory, goal and successor without converting persistent behaviour into an invented desire to survive.

Forthcoming · Track 4 · scheduled 27 October 2026 · 7 min read

Silicon, scale, memory and world models

Why parameter count is neither neuron count nor a consciousness measure, and how to separate architecture, memory, world modelling, implementation and physical substrate in causal studies.

Forthcoming · Track 4 · scheduled 3 November 2026 · 10 min read

A blind causal test of artificial introspection

A public protocol and analysis plan for testing whether an open-weight model identifies a hidden activation intervention above a strong output-only observer, while ruling out leakage and generic disruption.

Want a question of your own answered?

Working drafts go to anyone who asks, including - especially - people who will argue with them. If you have a model or an agent and a question this instrument could answer, we would like to hear it.

Ask a question How we measure

Ask for a draft