the lab
An illustrated mechanism

type your own message and press enter — the page follows

The journey of a prompt

What happens between pressing Enter and the first streamed word.

Watch, in slow motion, what really happens when you press Enter in a chatbot. The handful of words you type turn out to be only the last line of a much larger hidden document the machine reads before it replies.

Narrated by a synthetic voice — not a recording · ~0.3 MB

Scroll
The experiment

Your message is the last line of a much larger document; this lab lets you follow that document from Enter to the first streamed reply.

The surprise: you type about a dozen words. The machine receives around three thousand pieces of text. The other 99.5% is stitched on invisibly, every single turn — and it is a big part of what you sit and wait for.

For a long time I believed the obvious thing: I type a question, I press Enter, and my question goes to the chatbot. Twelve words, out the door. That picture is wrong in almost every detail. A chatbot remembers nothing between turns. So each time you hit Enter, the app quietly rebuilds the entire conversation — hidden instructions at the top, every earlier message in the middle, your new line at the very bottom — and sends the whole thing as one document. Your words are the last half-percent of it.

This page follows that document the whole way: what leaves, where it waits, how it is read, and why the reply comes back the way it does — a pause, then a stream.

What is illustrative

One piece of vocabulary you'll meet in a moment: the system doesn't count words, it counts tokens — short chunks of text, roughly one token for every three-quarters of a word (you'll see "tok" on the counters). The counts and clocks here are toy values, not measurements; real systems differ from one AI to another, and even by the hour of the day. The shape is what matters.

How to walk it: scroll down through the seven stops. At each one, drag a slider or press a button on the stage to your right and watch the document respond.How to walk it: scroll down through the seven stops. At each one, tap a lane, drag a slider or press a button on the stage above and watch the document respond.

You're reading the pocket version — everything works under a thumb, but on a desk the whole page answers your cursor. Worth a second visit.


01 · Enter

The part you see

A handful of words in a text box. Press Enter and they're gone — that was the whole event, as far as I ever thought about it.

The stage beside this holds the journey as the machine sees it. Right now it holds only what you typed — the message from the box at the top of the page. Watch how quickly that stops being the main character.

Try it — drag “your message” below the stage and watch the sliver breathe Your message is only the visible edge of the trip. Next: see what travels with your message ↓
02 · The envelope

Your message does not travel alone

The AI — the machine on the other end of the chatbot — keeps nothing between turns. Whatever it seems to remember about your conversation exists because the app quietly re-sends it — all of it — every single turn, stapled behind a system prompt you never wrote and will never see.

So the thing that actually leaves is a document: instructions at the top, the entire conversation in the middle, and your words at the very bottom. The last half-percent of it. The whole document travels every single turn, and every piece of it is counted.

Try it — drag the conversation longer, watch your share shrink The AI receives a document, not a lone message. Now turn that document into the units the system uses ↓
03 · Tokens

The document becomes numbers

Before the machine can read a word of it, the whole document is cut into tokens — sub-word pieces from a fixed menu. Your message isn't twelve words to the machine; it's fifteen or so numbered shards, sitting at the end of a few thousand others.

I've slowed tokenization down properly elsewhere. Here it matters for one reason: everything downstream — the queue, the clocks, the bill — is priced in tokens, not words.

Look for your message in the strip: the words have become numbered pieces. Go deeper: the tax on every other language → From here on, the system counts pieces, not words. The pieces are ready. Next they wait for a turn on the machine ↓
04 · The queue

You are not alone in the machine

Your document doesn't get a computer to itself. It lands in a queue beside strangers' documents — someone's homework, someone's contract, someone's six-thousand-token planning session — all of them sharing the same machine, handled together to keep it busy.

Some of the pause you feel before a reply isn't the machine doing sums at all. It's waiting for a seat.

Try it — hover the other lanes; everyone's document is longer than they thinkTry it — tap the other lanes; everyone's document is longer than they think Some of the pause is waiting for a seat. A seat opens. Now watch the document get read ↓
05 · Clock one — prefill

It reads everything before it says anything

Here's the part that rearranged my head. The AI can't glance at a document. It has to read all of it — system prompt, every old turn, your message — in one big pass, the whole thing at once, before it can begin. (This first pass has a name: prefill.) Only then does the first piece of the reply exist.

That pass is clock one: the pause before the first word. It grows with everything the AI must read — your fresh message included, though a long history usually dwarfs the line you just typed. Add the wait for a seat from the last station, and that's the pause you feel.

Try it — drag the conversation and watch which clock moves Clock one is the wait before the first token. The first token exists. Next, watch the rest arrive one piece at a time ↓
06 · Clock two — writing the reply

Then: one piece at a time

After prefill, the rhythm changes completely. The reply is generated one token at a time — each new piece appended to the document, then the machine decides the next. That steady beat is clock two: the gap between tokens, which you feel as the gap between words.

Two clocks, two very different rhythms: one reads everything in bulk, the other writes back one piece at a time. The pause and the stream you've felt in every chat are these two clocks, felt from the outside.

Try it — press “run the clocks” and watch the reply accrue at the modelled pace Clock two is the beat between generated pieces. Those pieces are generated. Now watch them arrive on your screen ↓
07 · The stream

The typing is (mostly) real

When the reply drips into your chat window, that isn't a loading animation. The AI really is deciding the reply one piece at a time, and the stream is those pieces reaching your screen — though the system may gather a few of them into each visible burst, and my demo groups them into whole words so you can read along.

And when you answer? The journey runs again from the top — with the reply you just watched folded into the document, making the next read a little heavier. The whole thing travels again, and every piece of it is counted again. That is why a long chat keeps getting slower: each turn re-reads everything before it. When speed or cost matters, start a fresh chat or trim the history.

Try it — press “stream it again” to re-watch the reply arrive The reply reaches you in pieces because it is produced in pieces. You have followed one turn. Send another, and the whole journey begins again ↓
·  one press of enter, slowed down
the document
system prompt · hidden
everything you've both said
the reply
system prompt620 tok
conversation · 24 turns2,473 tok
your message15 tok
the AI receives3,108 tok
the two clocksready
clock one · to the first token
the whole document, read first
clock two · between tokens
one token at a time

toy clocks — modelled, not measured

tap or hover a lane to read its size
what should I cook tonight with what's left in the fridge?

Next turn, this reply rides along too — ≈3,186 tokens, nothing forgotten, everything re-sent.

11 words
12 turns