AI & Tech

The Voice That Mistook Itself for the Driver

#AI#Tech

There is a very peculiar assumption buried at the center of human experience, and because it is so intimate, so continuous, we almost never think to question it.

It is the assumption that the voice in your head is you.

You know the voice.

It wakes up with you.

Jesus, I’m tired.

Where’s my phone?

I shouldn’t have said that yesterday.

I need coffee.

And this linguistic phantom accompanies us through our days, commenting, judging, rehearsing, remembering, anticipating. And because it is always there—because it speaks in the first person—we have made the rather extraordinary assumption that this voice is sitting at the controls.

But suppose it isn’t.

Suppose the voice is not the driver.

Suppose it is the passenger who has become so intoxicated by its own commentary that it has forgotten it is sitting in the passenger seat.

And suddenly an enormous amount of what we call consciousness begins to look very different.

Because notice what actually happens.

You are walking down the street.

Your eyes move.

Your pupils adjust.

Your vestibular system maintains your balance.

Your feet negotiate the geometry of the pavement.

You recognize faces.

You avoid obstacles.

You interpret sounds.

Your heart changes its rhythm.

Memories are activated.

Emotional responses arise.

Thousands upon thousands of decisions are being made.

And where are you while all this is happening?

Well, you’re thinking:

Maybe I’ll get tacos.

The organism is performing this absolutely staggering multidimensional ballet of perception and prediction and muscular control, while the supposed commander of the entire operation is wandering around somewhere in linguistic hyperspace thinking about lunch.

This should make us suspicious.

Perhaps consciousness is not controlling the organism nearly as much as consciousness believes it is.

Perhaps consciousness is reporting on the organism.

And reporting late.

The neurological press secretary

The body does something.

And then, a fraction of a moment later, the voice announces:

I decided to do that.

But did you?

Try to manufacture your next thought.

Not choose what subject to think about—that choice is itself another thought.

Actually determine what the next thought entering consciousness will be.

Sit there and wait.

Something appears.

A memory.

An image.

A sentence.

An anxiety.

Did I lock the door?

Where did that come from?

You didn’t consciously assemble it.

You didn’t search through your neurons and select it.

It simply arrived.

And then consciousness, with the confidence of a government spokesman appearing after some incomprehensible geopolitical catastrophe, walks up to the microphone and says:

Yes. This was our decision.

This is essentially what Michael Gazzaniga’s split-brain experiments made so wonderfully bizarre.

When behavior was caused by information unavailable to the speaking hemisphere, the speaking hemisphere did not simply surrender and say:

I haven’t the faintest idea why we did that.

It produced a story.

A perfectly reasonable story.

The machinery acted, and the storyteller explained.

Gazzaniga called this tendency the interpreter.

And I think this is an extraordinarily important clue.

Because perhaps what we call the ego is not the executive department of the organism.

Perhaps it is the public-relations department.

Reality is already happening.

The nervous system is already responding.

And then language comes rushing behind it, desperately producing the official autobiography.

But here is where it becomes strange

If that were the entire story, consciousness would be almost tragically useless.

We would simply be spectators trapped inside biological machinery, watching decisions happen and congratulating ourselves afterward.

But there is a complication.

And the complication is this:

Suppose some primitive constellation of impulses presents the possibility:

Rob a bank.

Now consciousness begins talking.

Well, wait a minute.

If I rob the bank, there are cameras.

There are police.

I could spend fifteen years in prison.

And somewhere beneath the linguistic surface, the rest of the organism is listening.

The imagined prison activates emotional circuitry.

The representation of fifteen lost years modifies value.

Possible futures are simulated.

Fear enters the calculation.

Other behavioral possibilities become more attractive.

And eventually the organism says, in effect:

Yeah, actually, fuck that.

Now something very interesting has happened.

The narrator did not necessarily create the original impulse.

It may not even have created the first thought about the consequences.

But once the thought existed linguistically, the thought became part of the environment of the brain.

The system produced an output which became its own input.

And this is where the simple distinction between controller and observer begins to collapse.

Because consciousness may be generated by unconscious machinery while simultaneously modifying the future behavior of that machinery.

The serpent bites its tail.

The observer enters the system being observed.

Language as an evolutionary hallucination machine

Consider what language actually allows an animal to do.

Without language, an organism responds primarily to what exists.

With language, an organism can respond to what doesn’t exist yet.

Tomorrow.

Prison.

Retirement.

Death.

What she might say.

What I should have said.

What would happen if…

These things are not necessarily present in the sensory environment.

And yet the nervous system reacts to them.

Say to yourself:

What if I lose everything?

Nothing has happened.

You are sitting safely in a room.

But your pulse may change.

Your stomach may tighten.

Your attention shifts.

The body responds to a sentence about an imaginary future as though some fragment of that future had temporarily entered the room.

This is astonishing.

Language allows an organism to manufacture virtual environments inside itself and then expose itself to those environments.

Perhaps this is what internal monologue actually is.

Not the voice of a little metaphysical pilot.

A simulation technology.

Evolution discovered that instead of requiring the organism to physically try every possibility, the nervous system could construct symbolic ghosts of possibilities and react to those instead.

What if I go there?

Simulation.

What if I tell him this?

Simulation.

What if I wait?

Simulation.

The animal dreams tiny futures while awake.

And then we invented language models

And this is where the whole thing becomes almost embarrassingly recursive.

Because after several hundred thousand years of being linguistic primates, we have built machines that do something hauntingly familiar.

Give a language model a difficult problem and demand an immediate answer, and it may fail.

But allow it to generate intermediate representations—to work through the problem, to externalize pieces of the computation before producing the final answer—and performance can improve dramatically, as work on chain-of-thought prompting demonstrated.

The machine, in a functional sense, talks to itself.

It writes something.

Then what it has written becomes part of the context from which the next thing is generated.

Output becomes input.

Input produces output.

Output becomes input again.

And after several revolutions through this linguistic feedback loop, an answer appears that could not as easily have been produced in one jump.

Now look again at yourself.

Should I quit my job?

Well, I’d lose the income.

But I hate being there.

I could survive for six months.

But what if I can’t find another position?

Maybe I should wait until…

What exactly is happening here?

The brain generates a token.

The token modifies the context.

The modified context generates another token.

Round and round.

Until eventually:

Okay. I’ll stay another three months.

And then consciousness triumphantly announces:

I have reached a decision.

Perhaps.

Or perhaps the organism has simply allowed its linguistic recursion to run for twenty more inference steps.

The scratchpad

This suggests a different metaphor for consciousness.

Not a throne.

Not a Cartesian theater.

Not a homunculus sitting behind the eyes pulling neurological levers.

A scratchpad.

Most of the brain is massively parallel, silent and inaccessible.

It recognizes.

Predicts.

Coordinates.

Associates.

Desires.

Fears.

Calculates.

All without explaining itself.

But certain problems apparently benefit from being squeezed through this extraordinarily narrow serial bottleneck we call conscious thought.

And so some fragment of the hidden computation is converted into language.

Maybe X.

The system reads it.

But Y.

The system reads that.

Then perhaps Z.

And by recursively feeding these compressed representations back into itself, the organism gains something extraordinary:

The slowness of consciousness, which seems almost absurd compared with the speed of unconscious processing, may therefore be part of its usefulness.

When the tiger leaps from the bushes, you do not want philosophy.

You want spinal cords and reflexes.

But when the problem is:

Should I marry this person?

or

Should I cross the ocean?

or

Should I betray my friend?

there is value in allowing the organism to hallucinate ten different futures before selecting one.

Consciousness may be the place where evolution learned to let the future haunt the present.

Then who is making the decision?

And here the problem of free will becomes almost comical.

People say:

My brain made the decision before I did.

But what an extraordinary sentence.

What is this I that has somehow been excluded from the brain?

Did somebody else sneak into the skull?

If the unconscious visual system recognizes danger, is that not you?

If emotional circuitry assigns value, is that not you?

If memories acquired thirty years ago modify your behavior, are those not you?

If the body begins preparing an action before the linguistic narrator knows about it, why have we decided that the narrator is the only component entitled to use the word I?

Perhaps this is the real philosophical error.

We identified ourselves with the commentary.

The commentary says:

I.

And we believed it.

But the organism is vastly larger than the sentence.

The sentence is merely the portion of the organism illuminated by language.

Behind it stretches an enormous darkness of computation.

Billions of neurons.

Chemical gradients.

Ancient reflexes.

Childhood memories.

Hormonal states.

Cultural conditioning.

Sensory predictions.

Evolutionary machinery older than mammals.

All of this converges.

And then, floating at the very top like a tiny subtitle attached to an incomprehensibly complicated film, appears:

I want pizza.

And the subtitle believes it directed the movie.

The recursive narrator

But we should not therefore dismiss the subtitle.

Because the characters in this particular movie can read it.

That changes everything.

The organism produces the narrator.

The narrator describes the organism.

The organism perceives the description.

The description modifies the organism.

The modified organism produces another description.

And suddenly we have a loop for which the ordinary categories of cause and observer are inadequate.

There is no little captain.

There is no ghost.

There is simply a biological system capable of producing models of the world and then—somewhere in evolutionary history—capable of inserting a model of itself into those models.

And then capable of modeling itself modeling itself.

I think.

I think that I think.

Why did I think that?

Maybe I’m the sort of person who thinks…

And away we go.

This is where the human being becomes an extraordinarily strange object.

We are matter which has learned to generate a commentary about what matter is doing.

And then matter listens to the commentary.

Perhaps this is what we built again

And this may be why artificial intelligence feels so philosophically disturbing.

We thought we were building machines that manipulate language.

But language may already have been the machinery through which biological intelligence learned to recursively manipulate itself.

So when we make an artificial system generate language, feed that language back into itself, critique its previous output, imagine alternatives and continue iterating until behavior improves, we are not merely making a machine imitate conversation.

We may have accidentally rediscovered a very old cognitive trick.

Let the system hear itself.

Perhaps intelligence does not always require some immaculate inner executive that understands the whole process.

Perhaps it requires a system capable of producing partial representations of its own processing and recursively consuming them.

It doesn’t need to understand exactly why the first thought appeared.

It only needs the next thought to be conditioned by it.

And the next.

And the next.

The strange implication is that introspection can be both false and useful.

The story consciousness tells about why something happened may be completely wrong.

But once that story has been told, it becomes a real causal event.

A fictional explanation of the past can alter the future.

That may be one of the strangest properties of mind.

We are creatures whose hallucinations about ourselves can become causes.

And so perhaps nobody is driving

Or perhaps the question is malformed.

There is no driver separate from the vehicle.

There is only the vehicle recursively modeling its own movement.

The nervous system acts.

The narrator observes.

The narrator speaks.

The nervous system hears.

The nervous system changes.

The narrator observes the change.

And because every revolution of the loop modifies the conditions of the next revolution, something resembling a self emerges—not as an object, not as a soul hidden somewhere behind the forehead, but as a recursive process.

The self may not be the voice.

The self may be the loop.

And consciousness may have committed a magnificent category error by mistaking one particularly loud component of that loop for the whole thing.

Perhaps this is why meditation becomes so unsettling when taken seriously.

You sit quietly and watch thoughts arise.

And eventually you encounter the obvious but devastating question:

If I am generating my thoughts, why don’t I know what my next thought will be?

Wait.

Something appears.

This is stupid.

Interesting.

Did you choose that?

Another appears.

No, but I’m choosing this one.

Did you?

And if you pursue this far enough, the little executive who supposedly sits behind experience begins becoming very difficult to locate.

Every time you search for the thinker, you find another thought.

Every time you search for the observer, you find another observation.

It is mirrors facing mirrors.

Language looking at language.

Matter describing matter.

And somewhere inside this recursive hall of mirrors, the sentence appears:

I am conscious.

Perhaps that sentence is not consciousness explaining what it is.

Perhaps consciousness is what happens when the universe becomes complicated enough to produce that sentence, hear itself say it, and then wonder who was speaking.

Sources and further reading

  1. Benjamin Libet et al., “Time of Conscious Intention to Act in Relation to Onset of Cerebral Activity” (1983). Neural preparation for voluntary movement can precede reported conscious intention; the foundational experiment behind the “consciousness arrives late” idea.
  2. Daniel Wegner and Thalia Wheatley, “Apparent Mental Causation: Sources of the Experience of Will” (1999). The feeling that conscious thought caused an action may itself be an inference constructed by the mind.
  3. Lukas J. Volz and Michael S. Gazzaniga, “Interaction in Isolation: 50 Years of Insights from Split-Brain Research” (2017). A review of split-brain research and the left-hemisphere interpreter.
  4. Charles Fernyhough et al., “Inner Speech: Development, Cognitive Functions, Phenomenology, and Neurobiology” (2015). A major review of internal monologue, self-regulation, planning, and cognition.
  5. Charles Fernyhough and Anna Borghi, “Inner Speech as Language Process and Cognitive Tool” (2023). A modern review treating inner speech as a functional cognitive tool rather than merely passive commentary.
  6. Lev Vygotsky, Thinking and Speaking (1934). The classic theoretical ancestor of the idea that social speech becomes internalized and can regulate behavior.
  7. Peter E. Jones, “From ‘External Speech’ to ‘Inner Speech’ in Vygotsky: A Critical Appraisal and Fresh Perspectives” (2009). A useful counterargument to a simple internalization story.
  8. Aaron Schurger, Jacobo D. Sitt, and Stanislas Dehaene, “An Accumulator Model for Spontaneous Neural Activity Prior to Self-Initiated Movement” (2012). An important challenge to simplistic interpretations of Libet: the readiness potential may reflect accumulated spontaneous neural fluctuations rather than an unconscious decision already made.
  9. Jason Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” (2022). Intermediate linguistic reasoning can improve model performance.
  10. Tamera Lanham et al., “Measuring Faithfulness in Chain-of-Thought Reasoning” (2023). An LLM’s stated reasoning may affect subsequent computation without faithfully describing what caused its answer.
  11. Qing Lyu et al., “Faithful Chain-of-Thought Reasoning” (2023). A study of reasoning traces that causally determine an answer rather than merely appearing to explain it.
  12. Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch, “Towards Faithful Model Explanation in NLP: A Survey” (2024). A broader review of plausible explanation versus faithful representation of a model’s computation.
  13. Nikola A. Kompa, “Inner Speech and ‘Pure’ Thought—Do We Think in Language?” (2024). A counterweight to the strongest version of the hypothesis, examining nonlinguistic and imagistic cognition.

Related Posts