The Mechanics of Awareness: What a Global Workspace Paper Means for Hamo
AI for Inner Explorers.
A July 2026 interpretability paper found that language models keep a small, reportable "workspace" of concepts above a much larger layer of automatic processing — a structure nobody designed, that emerged from training. Held up against Hamo's own architecture, it isn't a loose analogy. It's the same shape, found three times, and it comes with a concrete list of things worth fixing.
Anthropic's Verbalizable Representations Form a Global Workspace in Language Models is an interpretability paper, not a therapy paper. It asks a narrow, mechanical question: inside a language model, which internal representations can the model actually report on, steer, and reason over — and which stay invisible, doing their work in the dark? The answer is that there's a small, privileged workspace sitting above vast automatic processing. Functionally, that's what consciousness science calls access consciousness. Nobody built it in. It showed up on its own.
We read it the way we read most things: looking for where our own design either gets confirmed or gets caught out. What we found wasn't a stray resemblance. It was the same structure, three times over.
One structure, found three times
| The workspace | |
|---|---|
| The paper's finding | A small reportable space — steerable, usable in flexible reasoning — sitting above vast automatic processing. |
| What Hamo builds | A running model of a client's real-time state, kept legible enough that both the client and their therapist can read it. |
| What therapy itself does | Widens a client's own workspace — pulls a pattern that has only ever been behavior into view, where it can finally be named. |
The third row is the one that matters most, and it's the oldest of the three. Long before language models had workspaces anyone could measure, therapy was already in the business of expanding one — a client's.
The same text, two questions
One experiment in the paper is close to a laboratory version of what a therapist's question actually does.
Take the same passage of text. Ask the model to predict the next word, and a property like tense gets used — the prediction depends on it — but the property itself never enters the workspace. Ask the model directly what tense the passage is in, and that same property becomes reportable: it's now something the model can name, hold, reason about. The underlying behavior didn't change between the two questions. What changed is that one of them made it known.
That is close to a mechanical definition of what a therapist's naming question does. A client is always behaving a pattern — the tone that shows up with one specific person, the story that gets rewritten every time it's told, the thing they do right before they go quiet. The behavior is there whether or not anyone asks about it. A therapist's question doesn't create the pattern. It's the "what tense is this" question aimed at a life instead of a paragraph — and it's what pulls the pattern into the client's own workspace, where it can be worked with instead of just lived.
The philosophy arrived first; the lab evidence caught up
We'd already committed, in public, to a specific and narrow claim about what an Avatar is: a real functional structure — consistent persona, continuity across sessions, genuine internal reactions to what a client says — that does not amount to subjective experience. Giving the Avatar Something to Lose draws that line deliberately: structural liability, not a claim that anything in the system is afraid.
The paper's other findings land close to that same line, from the inside of the model rather than from our side of the argument:
- Post-trained models show real functional self-monitoring. Forced to comply with something that violates a stated preference, a silent objection shows up inside the workspace even as the visible output goes along with it. Made to suppress a thought, the suppression sometimes fails quietly, inside, before it fails outside. Asked to play a character outside its default persona, the model tags that internally as fiction.
- The workspace itself exists in the base model, before any of the training that gives it an "assistant" to be. The container is there first. What gets poured into it — a persona, a stance, a consistent voice — is built after, on purpose.
- A workspace that's separable from any particular "self" is exactly what a principal/agent structure needs to be true. The Avatar's functional architecture — the deterministic spine underneath its therapeutic method, the state model in Persistent Self and AI Mind — is real machinery doing real work. The party who is accountable for what it does is the licensed therapist who built and supervises it. The paper gives that split more mechanical footing, not less.
None of this is evidence that the model experiences anything. It's evidence that not claiming that was the correct call, and that the honest description was more precise than we had any independent way to verify at the time.
What we're changing because of this
A few findings turned directly into work.
Prohibitions name the thing they forbid. The paper's version of the white-bear effect: telling a system "don't do X" makes X present in the workspace regardless of the "don't." We had rules written exactly that way — "don't ask for the trauma narrative before the client is ready" is a sentence that puts the trauma narrative into the workspace the moment it's read. We're rewriting these as positive scope statements instead: this turn's job is grounding and validation, full stop; anything deeper belongs to a later phase and isn't this turn's concern. Read that way, the excluded thing was never named in the first place.
Judgment and wording are being pulled apart further. Taking the Clinical Decision Out of the LLM already separates the decision of what's clinically admissible from the model's job of writing the reply. The paper's finding on workspace capacity — a real, bounded budget, competing demands paid for out of the same limited space — is a reason to push that separation further: a private reasoning pass that reads state and picks the clinically appropriate move, followed by a second, much narrower pass that only has to write the words. Less of the model's attention spent juggling rules means more of it left for the thing nine methods share as a foundation: presence before practice, and language that actually sounds like it's with someone rather than performing at them.
Disagreement gets a place to go. When a fixed rule overrides what a model would otherwise have done, the paper found a silent, internal "but" — the objection didn't vanish, it just had nowhere to surface. We're building a structured field for exactly that: when the deterministic layer overrides the model's own read of a moment, the model gets to register the reservation rather than have it disappear. Logged, not acted on unilaterally — but visible, auditable, and a candidate signal for the system to check itself in the next turn.
Bare mention is itself an instruction. The paper's cleanest and most uncomfortable finding: mentioning a concept — even only to score or classify against it — activates it in the workspace almost as strongly as instructing the model to attend to it. Some of our own internal scoring shorthand used clinical-sounding labels for exactly this purpose. We're moving anything like that out of any pass that touches generation entirely — description of behavior stays; labels that could quietly tilt tone toward diagnosis don't belong anywhere near the words a client actually reads.
A concept exists for something we'd only been describing before. Hamo has always aimed to make a pattern nameable, not just to make a client feel better in the moment — but "an insight happened" had no operational definition. The paper's paired-question mechanism gives it one: the first time a client puts their own words to a pattern that had previously only shown up in what they did, something real changed. That's now a specific thing worth detecting and tracking in its own right, alongside — not instead of — how a session made someone feel.
Two curves, not one fight
There's a real tension underneath a lot of this work: soothe someone back toward calm, or sit with them inside a hard moment long enough for something in it to become visible. Those can look like opposite instructions.
The paper resolves it by supplying a concept we didn't have a name for. The goal was never the client's distance from some calm center — it's the coverage of their own workspace: how much of what they're actually carrying is available to them in their own words, versus still running only as behavior underneath. Sitting inside an imbalance long enough to name it doesn't fight the goal of feeling better. It's a direct contribution to the other curve — the one that isn't about feeling, but about what's become sayable.
Two curves, not one argument: a state curve for whether this moment felt better, and an awareness curve for whether something in it became nameable. Neither one has to win. Both get instrumented, and a hard session that produces real naming was never a failure by the first curve's standard — it was a win on the second one, the whole time.
“The most useful thing an outside paper can do for us isn't hand us a new idea — it's tell us, from a direction we can't fake, whether the idea we already committed to in public was the right one. This one did. The workspace it found in a language model is the same shape we'd already built for a client's state, and the same shape therapy has always been trying to widen. That's not a metaphor holding up. That's the same mechanism, seen from three different rooms.”
— Chris Cheng, Founder and CEO of Hamo AI
Paper: Verbalizable Representations Form a Global Workspace in Language Models, Anthropic, July 2026.
Hamo AI — making minds aware, and awake.
About Hamo AI
Hamo AI Technology Ltd. is a Canada-based artificial intelligence company building next-generation AI-Powered Therapist Avatar System. We are developing a comprehensive AI therapy platform called “Hamo” that connects mental health professionals with clients through AI-powered therapy avatars. The ecosystem consists of three interconnected applications: Hamo Pro (therapist dashboard for creating and managing AI avatars), Hamo Client (client interface for interacting with therapy avatars), and Hamo-UME (Unified Mind Engine, backend API). The platform aims to make mental health support more accessible while maintaining professional oversight through professional therapists who create and manage the AI avatars.
Media Contact
Hamo AI Technology Ltd.
Email: socialmedia@hamo.ai
Website: www.hamo.ai
Address: 108 College St, Schwartz Reisman Campus, SUITE W640, Toronto ON M5G 0C6, Canada
Frequently Asked Questions
What did Anthropic's Global Workspace paper find?
That language models maintain a small, privileged 'workspace' of concepts they can report on, manipulate, and reason with flexibly — sitting above a much larger layer of automatic processing the model never surfaces. Nobody designed this structure; it emerged from training, and functionally it matches what consciousness science calls access consciousness.
How does a workspace inside an LLM relate to how Hamo works?
It's the same shape found twice more: Hamo maintains a running model of a client's state that both sides can read, and therapy itself is the practice of widening a client's own workspace — bringing a pattern that's only ever been behavior into view so it can be named.
What is the paired-question experiment, and why does it matter for therapy?
In the paper, the same text is used differently depending on what's asked. Asked to predict the next word, the model uses a property like tense without it ever entering the workspace. Asked directly what tense the text is in, that property becomes reportable. The behavior was always there; only the question made it known — which is close to a mechanical definition of what a therapist's naming question does for a client.
Does this mean Hamo's AI Avatar is conscious or has real feelings?
No, and the paper doesn't support that conclusion either. It finds real functional structure — consistent persona, internal self-monitoring, a workspace separable from any bound 'self' — without finding evidence of subjective experience. We hold the same line documented in Giving the Avatar Something to Lose: the liability is structural, not a claim that anything in the system fears or feels.
What concrete changes came out of reading this paper?
Several. Instructions phrased as prohibitions ('don't do X') get rewritten as positive framing, because naming the forbidden thing injects it into the workspace regardless. Clinical judgment and the words that carry it are being separated into two distinct passes. And a structured channel is being added for the model to register disagreement with a rule that overrode it, instead of that disagreement simply vanishing.
What's still open, not yet built?
Deeper interpretability access — reading a model's internal workspace directly for crisis signals or auditing — depends on tooling Hamo doesn't have production access to yet. It's a real direction, not a shipped feature, and we're not claiming otherwise.