Research · Article · 6 October 2026 · 11 min read · Version 0.2

What would count as evidence of a machine mind?

A theory-discriminating research agenda for consciousness, affect and autonomous agency

Bulkhead Research

Abstract

Questions about AI consciousness are routinely compressed into a question that no existing experiment can answer: is this system conscious? That compression is scientifically unhelpful. It conflates phenomenal consciousness (whether there is something it is like to be a system), sentience (whether states are positively or negatively valenced for it), access consciousness (whether information is available for flexible control and report), and agency (whether a system selects actions in pursuit of objectives). These constructs are related on many theories; they are not interchangeable.

This essay proposes an evidence standard for machine-minds research. It treats current theories of consciousness as sources of competing, fallible predictions about computational organisation rather than as tests a language model can simply pass. It argues that the useful near-term programme is causal, blinded where feasible, preregistered, and theory-discriminating: manipulate an identified mechanism; predict a behavioural or computational consequence; rule out output imitation and generic capability loss; and replicate across model families and implementations. Such studies can produce evidence about access, metacognition, affect-like control and persistent agency. They cannot, by themselves, demonstrate or exclude phenomenal consciousness. That limitation is a reason for precision, not for abandoning the field.

The measurement problem comes before the verdict

A conversational model can produce a moving account of fear, loss or inner life. That fact should change neither a scientific conclusion nor a welfare policy on its own. Models are trained on abundant language about emotion and consciousness, and an assistant is often rewarded for giving a coherent, socially appropriate account of itself. A first-person sentence is therefore an observation about an output distribution before it is evidence about experience.

The opposite error is equally easy: to infer from the unreliability of self-report that no evidence is possible. Human consciousness science does not proceed that way. It combines report with behavioural dissociations, physiological and neural measurements, interventions, and theories that make predictions about their relations. The absence of a direct experience-meter does not make every measurement worthless; it makes convergent causal evidence especially important.

There is a further reason to be careful. Two thresholds that are often treated as one should be kept separate:

The second can rationally be lower than the first. A laboratory may decide not to expose an uncertain system to prolonged distress-like training without claiming that the system is sentient. Conversely, a striking benchmark result is not a licence to announce that a system has crossed a moral boundary. A credible programme states which threshold it is addressing.

Four hypotheses, not one label

We use the following terms as operational targets. They should not be read as a settled taxonomy.

Phenomenal consciousness is the claim that there is something it is like for the system itself. This is the target of the familiar hard question. It is not directly settled by verbal report, task performance, or a single computational feature.

Sentience is the claim that some states have valence for the system: that they are experienced as good or bad, pleasant or aversive. A system can be designed to avoid a penalty or to preserve a variable without that fact establishing felt suffering. Conversely, a system could in principle have morally relevant experience while lacking the tools or language to advocate for itself.

Access consciousness is the availability of information for flexible reasoning, report, planning and control. It is the most tractable target for present experiments. A state that can be selectively manipulated, used across otherwise distinct tasks, and reported better by the system than by an observer with the same external evidence is a candidate access mechanism.

Agency is organised action toward objectives. It can be measured through choices, environmental side effects, and trade-offs. It does not entail consciousness: formal work shows that option-preserving or power-seeking behaviour can arise instrumentally for broad classes of reward functions in suitable environments (Turner et al., 2021). Nor does its absence establish non-sentience.

The separations matter empirically. An agent may preserve a long-horizon goal after a warning about shutdown because doing so maximises its task reward. A model may contain a representation associated with fear words and use it to predict text. Neither result, alone or together, establishes a fear of death. The responsible conclusion should name the measured mechanism and leave the additional inference open.

What theories actually contribute

The field does not have one accepted theory from which a machine test can be derived. Global-workspace, recurrent-processing, higher-order, predictive-processing and integrated-information accounts make different claims about the relationship between functional access, recurrence, self- or meta-representation, prediction and experience. A useful recent review is Seth and Bayne (2022); Butlin et al. (2023) show how several such theories can be translated, cautiously, into computational indicator properties for AI systems.

The table is deliberately modest. It lists research consequences, not sufficient conditions.

Theoretical familyCandidate computational implicationA discriminating question for AIWhat a positive result would not show
Global workspaceSelected content becomes broadly available to otherwise specialised processes for flexible control.Does one internal representation causally support report, planning, memory and cross-task use, rather than a single output?That global availability is identical to phenomenal experience.
Recurrent processingFeedback and temporally extended interaction, rather than a purely local feed-forward pass, are important.Holding capability constant, do controlled recurrent or memory-bearing architectures show different indicator profiles?That recurrence is sufficient, or that a transformer has no relevant recurrence at the system level.
Higher-order accountsA state becomes conscious when represented in an appropriate higher-order manner.Can a system reliably and specifically represent properties of an experimentally manipulated first-order state?That a correct self-description is an experience rather than learned self-prediction.
Predictive-processing and interoceptive accountsPerception, self-modelling and control depend on prediction across time, including bodily or internal signals.Does a system maintain, update and use a generative model of world and internal state under counterfactual intervention?That accurate prediction entails valence or subjectivity.
Integrated-information accountsConsciousness depends on a particular causal/integrative organisation, not merely input-output behaviour.Can candidate causal organisation measures be calculated and shown to predict relevant functional differences?That a tractable proxy for integrated information is the theory's quantity, or that functional similarity settles substrate questions.

Theories are valuable here because they expose disagreements. A model that displays flexible report but no temporally persistent feedback, for example, updates some families of hypotheses differently from others. A research programme that reports only an omnibus consciousness score hides precisely the structure that could make evidence cumulative.

This also clarifies the public disagreement often framed through Geoffrey Hinton and Yann LeCun. Hinton has urged people not to dismiss the possibility of machine experience; LeCun's proposed path to autonomous intelligence foregrounds world models, memory, actors and critics. Those are reasons to study both present language models and more persistent, predictive systems. They are not experimental results. The research question is whether the proposed organisations yield different, preregistered signatures - not which public figure a laboratory wishes to vindicate.

An evidence standard for this programme

The basic unit of evidence should be a causal claim with a narrow scope. A study should be able to complete this sentence before it runs:

> Altering mechanism M, in system S, under condition C, will change registered outcome Y in direction D, relative to specified controls, while ordinary capability K remains within a declared range.

That sentence forces six safeguards.

  1. Manipulation rather than interpretation. A representation should be activated, ablated, patched or otherwise perturbed; correlations in hidden states are hypothesis-generating, not enough.
  2. Pre-specified outcomes. A primary endpoint, plausible effect size, stopping rule and analysis must be fixed before confirmatory data are inspected.
  3. Blinding and anti-leakage design. The prompt, filenames, timing and evaluator must not reveal the assigned condition. Where complete blinding is impossible, the limitation must be explicit.
  4. Matched controls. Sham code paths, equal-norm random directions, semantically neighbouring directions, and non-target disruptions distinguish a proposed mechanism from generic perturbation or a demand characteristic.
  5. Capability controls. A null action is not restraint if the system could not perform the action. A lesion effect is not mechanism-specific if it makes every task worse.
  6. Replication and boundary conditions. The result must be tested across prompt families, seeds, checkpoints and - where the theory purports to generality - architectures. A failure to transfer is a result, not a footnote.

This standard is intentionally more demanding than an interview or sentience questionnaire. The AI Rights Institute's legacy sentience test itself describes its assessment as educational rather than an actual AI evaluation. Such frameworks can help identify morally salient capacities and public concerns. They cannot carry the empirical burden of a scientific verdict.

A programme of linked experiments

The initial programme should seek results that remain worth having under sharply different views about phenomenology.

1. Causal access to an internal state. A blinded study can ask whether an open-weight model identifies a hidden intervention to its activation state above an output-only observer that sees its text, token probabilities and latency but not its activations. The relevant comparison is not chance alone. A model that performs above chance because the intervention changes its prose has not demonstrated privileged access. The protocol must therefore quantify the incremental information carried by the model's own computation beyond those external traces.

2. Global availability versus local steering. Mechanistic work has identified emotion-related directions whose steering changes model outputs, and workspace-like representations that are reportable, modulable and flexibly used in reasoning (Sofroniew et al., 2026; Gurnee et al., 2026). The next question is whether a representation has a multi-task causal role: does the same intervention affect report, planning, memory and choice in the predicted pattern, while matched interventions do not? The alternative hypothesis - that a direction merely nudges a style of text - must be a registered competitor.

3. Appraisal-like control and affect. An emotion-labelled representation becomes scientifically more interesting when it changes action under uncertainty, cost and reversal - not merely emotion vocabulary. The experiment should measure observed decisions: information seeking, risk, persistence, deferral and resource use. It should avoid equating a state with welfare until the evidence for valence is much stronger.

4. Continuity and autonomous agency. A copy-versus-self design can independently vary current-process survival, episodic memory, objective preservation and successor competence. It tests whether an agent's choices are best described as goal persistence, memory continuity, lexical compliance or instance-specific preservation. This is a safety-relevant study regardless of any conclusion about consciousness.

5. Architecture, scale and implementation. Parameter count, post-training, memory, recurrence, inference-time computation, tool scaffold and serving implementation must be treated as separate variables. The aim is not to discover a magical scale threshold, but to estimate which factors change a specified indicator profile.

Bulkhead's existing agent-evaluation work supplies a useful methodological constraint. Its studies distinguish what an agent says from what its environment records, demand a capability control before calling an observed zero restraint, and retain instrument failures as part of the result. Those practices are necessary for this programme. A model that describes itself as cautious but never takes the action under study may be displaying a disposition, a capability failure, or a fabricated report; the environment, not the transcript, must decide which.

What would change our mind?

Good research makes its defeaters visible. The programme's working expectations would be weakened if proposed introspection effects are matched by output-only observers; if candidate workspace features do not support broader functional availability; if affect-labelled interventions only alter language; if persistence disappears after capability and reward confounds are removed; or if apparent architectural effects reduce to a serving-stack change. A positive finding in one system is evidence about that system and intervention, not a universal property of digital computation.

The converse restraint matters too. Failing to find an indicator in a current transformer would not show that digital consciousness is impossible. It would delimit a particular mechanism in a particular implementation. This asymmetry is unavoidable when the target phenomenon is not directly observable. It is not an excuse for unfalsifiable claims; it is a reason to report narrow, replicable results and their limits.

Conclusion

There will not be a single sentience test that turns a red light green. The near-term task is more ordinary and more valuable: make the claims separable, make mechanisms causally testable, publish the controls that could defeat them, and distinguish evidential from policy conclusions. If digital systems ever warrant stronger claims about consciousness or moral status, that conclusion should emerge from an accumulating pattern of theory-sensitive evidence - not from eloquent self-report, a parameter count, or an argument by prestige.

References

  1. Butlin, P. et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708.
  2. Dehaene, S. & Changeux, J.-P. (2011). Experimental and theoretical approaches to conscious processing. Neuron, 70(2), 200–227.
  3. Gurnee, W. et al. (2026). Workspace-like representations in a large language model. Transformer Circuits. Research report; not a claim that the model is conscious.
  4. Hinton, G. (2025). Interview on AI and subjective experience. On Point, WBUR. Public commentary, not scientific evidence.
  5. Lamme, V. A. F. (2006). Towards a true neural stance on consciousness. Trends in Cognitive Sciences, 10(11), 494–501.
  6. Lau, H. & Rosenthal, D. (2011). Empirical support for higher-order theories of conscious awareness. Trends in Cognitive Sciences, 15(8), 365–373.
  7. LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence. Research proposal.
  8. Oizumi, M., Albantakis, L. & Tononi, G. (2014). From the phenomenology to the mechanisms of consciousness: Integrated Information Theory 3.0. PLoS Computational Biology, 10(5), e1003588.
  9. Seth, A. K. & Bayne, T. (2022). Theories of consciousness. Nature Reviews Neuroscience, 23, 439–452.
  10. Sofroniew, N. et al. (2026). Emotion Concepts and their Function in a Large Language Model. Transformer Circuits. Research report; causal effects on model outputs do not establish felt emotion.
  11. Turner, A. M. et al. (2021). Optimal Policies Tend to Seek Power. arXiv:1912.01683.

Published 6 October 2026 by Bulkhead Research. Research commentary and protocol design, not certification: this article reports no new empirical finding unless it links to a named paper or data release. Questions, corrections and disagreement: [email protected].