Being a candid mechanical account of the Large Language Model — what truly happens between your question and its answer — rendered in brass, in the manner of Mr. Charles Babbage.
Everything below actually runs: the gauges are moved by a real (if miniature) engine of probability computed live on this page. The plaques state, candidly, how the full-sized article differs — and terms under a dotted line yield their modern meaning to the pointer.
Language Engine
To operate: 1 — write a few opening words upon the card, or take a ready-punched card from the rack; 2 — pull Begin the Demonstration and follow the works, pressing Proceed at each station; 3 — thereafter Turn the Crank for each further word, or let the engine Run. Mind the boiler.
Your words are punched upon a card at the clavier — a stroke to the column — and fed to the reader. The engine does not hear you — it receives marks.
In truth — a language model gets only text. It has no voice to hear, no eyes to see the world, and no idea who is asking — except what the words themselves give away. Everything it does next comes from those letters, and from how its parts are set.
Why a few words and not a question? A chat engine seems to answer questions, but underneath it only ever continues text. Behind the scenes, the conversation is written out like a script — Operator: … Engine: — — and the machine carries the script on from where it stops. The next thing in the script is always the engine’s reply, so it looks like an answer. The rules for how it should behave (the system prompt) are just more text, put in front of yours. This cabinet skips the script and asks for opening words directly. The punch above the slot is real history, too: cards like these were cut one column at a time on a hand punch. Someone always has to turn words into holes.
The card is read and matched to slugs — standard pieces from the engine’s type-case. Each slug bears a catalogue number, and from here onward the engine handles numbers only.
In truth — the text is chopped into tokens — pieces from a fixed list of about 50,000 to 200,000. A common word is one token. A rare word is cut into parts (un·fathom·able). Each piece has a number, and the number is all the machine ever sees. It never sees letters or words the way you do.
This cabinet’s catalogue holds — slugs; a slug marked ✶ was cut fresh and is a stranger to it. The rack holds only 24 slugs — its context window. Push in more and the oldest fall off the far end, forgotten. (A modern engine’s rack holds 200,000 or more.)
In truth — this drawer is the model’s entire vocabulary. Every reply it will ever compose is set from these sorts and no others; a word absent from the drawer cannot be said. A real engine’s case holds 50,000–200,000 sorts — larger, but every bit as closed. Sorts marked ✶ were cut fresh this session; sorts aglow are out at the composing line.
Each catalogue number is looked up in the Great Ledger, and the slug is exchanged for a column of dial-settings — its position upon the map of meaning.
In truth — each token is turned into an embedding: a long list of numbers. This cabinet shows eight dials; a big engine uses more than ten thousand. Nobody wrote those numbers by hand. The machine learned them, and it learned to give words that are used alike numbers that are alike. To the machine, a word’s meaning is just where it sits among the others.
Here is the celebrated mechanism. Every slug is joined to every earlier slug by an adjustable linkage; the engine draws some taut and lets others hang slack. A word discovers what it means by choosing what to regard.
In truth — this is attention, the invention (2017) that makes these engines possible. For each position the model computes how much every earlier token should bear upon it — how it finds its noun, how bank leans toward the river or the counting-house. Dozens of heads work in parallel, each minding a different sort of relation: shown here, a nickel head that favours the recent, a copper head that favours old acquaintances.
This is a sketch. In the real engine these links are worked out fresh at every floor, from the numbers on the dials.
The columns now descend through the Mill — floor after identical floor of linkage-and-gearwork, each refining the settings handed down by the last.
In truth — attention plus a small calculator makes one layer. Layers are stacked — dozens of them, close to a hundred in a big model. A word’s numbers are changed a little at every floor. Low floors seem to handle spelling and grammar; higher floors handle meaning and intent — or so the people who study these machines report. The adjustable parts are called parameters. A big model has hundreds of billions of them, and not one was set by hand. In training, the engine read a huge library and guessed each next word. Every time it guessed wrong, every screw was turned a hair. Do that a great many times and that is the whole of its schooling. A finishing school comes after — training on what people prefer — to make it helpful, honest, and polite.
This cabinet, for comparison, owns — counting-wheels, learned from a library of — sentences.
At the mill’s outfall the engine renders its judgment — not a word, but a pressure upon every gauge in the house: a score for every slug in the catalogue at once.
In truth — at every step the machine gives a score to every token it knows — all 50,000-odd of them, most tiny. It does not pick the next word here. It rates them all. These gauges are real: they show this engine’s actual arithmetic for the words on the rack right now, and they answer to the boiler lever at Station VII if you move it.
For the first two words after a new card, the stop-mark and the full stops are held at zero, so a one-word card cannot be answered with silence. From the third word on, the engine may stop when it likes.
One slug must be drawn. The lots are cast in proportion to the gauges — and the boiler decides how strictly. Run cold, the favourite is all but certain; run feverish, and outsiders take their chance.
In truth — the next word is chosen by a weighted lottery — the trade word is sampling. Temperature changes the odds exactly as the drum shows. Turn it down and the favourite almost always wins: safe, but it repeats itself. Turn it up and the odds flatten: more surprising, and at the far end, nonsense. That is why the same question, asked twice, can get two different answers. Try the lever — the gauges change at once.
The chosen slug is stamped upon the tape — and carried straight back to the rack, where it joins your question. The whole engine then turns over again, from card to lottery, to choose the word after. One revolution, one word.
In truth — a language model writes autoregressively: one token per pass, each pass re-reading all that came before. There is no finished answer waiting inside; each word is minted the moment before you read it. And the engine stops as it speaks — by predicting a special stop token (printed here as ∎), whereupon it rests.
A real engine keeps notes (a cache) so it need not redo old sums on every turn. But in principle, every turn reads the whole rack again from the start.