Being a candid mechanical account of the Large Language Model — rendered in brass, in the manner of Mr. Charles Babbage, and schooled upon nothing whatever but the Sonnets of Mr. William Shakespeare.
Everything below actually runs: the gauges are moved by a real (if miniature) engine of probability computed live on this page. The plaques state, candidly, how the full-sized article differs — and terms under a dotted line yield their modern meaning to the pointer.
Language Engine
To operate: 1 — set the schooling above; 2 — write a few opening words upon the card, or take a ready-punched card from the rack; 3 — pull Begin the Demonstration and follow the works, pressing Proceed at each station; 4 — thereafter Turn the Crank for each further word, or let the engine Run. Mind the boiler.
Your words are punched upon a card at the clavier — a stroke to the column — and fed to the reader. The engine does not hear you — it receives marks.
In truth — a language model gets only text. It has no voice to hear, no eyes to see the world, and no idea who is asking — except what the words themselves give away. Everything it does next comes from those letters, and from how its parts are set.
Why a few words and not a question? A chat engine seems to answer questions, but underneath it only ever continues text. Behind the scenes, the conversation is written out like a script — Operator: … Engine: — — and the machine carries the script on from where it stops. The next thing in the script is always the engine’s reply, so it looks like an answer. The rules for how it should behave (the system prompt) are just more text, put in front of yours. This cabinet skips the script and asks for opening words directly. The punch above the slot is real history, too: cards like these were cut one column at a time on a hand punch. Someone always has to turn words into holes.
The card is read and matched to slugs — standard pieces from the engine’s type-case. Each slug bears a catalogue number, and from here onward the engine handles numbers only.
In truth — the text is chopped into tokens — pieces from a fixed list of about 50,000 to 200,000. A common word is one token. A rare word is cut into parts (un·fathom·able). Each piece has a number, and the number is all the machine ever sees. It never sees letters or words the way you do.
This cabinet’s catalogue holds — slugs; a slug marked ✶ was cut fresh and is a stranger to it. The rack holds only 24 slugs — its context window. Push in more and the oldest fall off the far end, forgotten. (A modern engine’s rack holds 200,000 or more.)
In truth — this drawer is the model’s entire vocabulary. Every reply it will ever compose is set from these sorts and no others; a word absent from the drawer cannot be said. A real engine’s case holds 50,000–200,000 sorts — larger, but every bit as closed. Sorts marked ✶ were cut fresh this session; sorts aglow are out at the composing line.
Each catalogue number is looked up in the Great Ledger, and the slug is exchanged for a column of dial-settings — its position upon the map of meaning.
In truth — each token is turned into an embedding: a long list of numbers. This cabinet shows eight dials; a big engine uses more than ten thousand. Nobody wrote those numbers by hand. The machine learned them, and it learned to give words that are used alike numbers that are alike. To the machine, a word’s meaning is just where it sits among the others.
Here is the celebrated mechanism. Every slug is joined to every earlier slug by an adjustable linkage; the engine draws some taut and lets others hang slack. A word discovers what it means by choosing what to regard.
In truth — this is attention, the invention (2017) that makes these engines possible. For each position the model computes how much every earlier token should bear upon it — how it finds its noun, how bank leans toward the river or the counting-house. Dozens of heads work in parallel, each minding a different sort of relation: shown here, a nickel head that favours the recent, a copper head that favours old acquaintances.
This is a sketch. In the real engine these links are worked out fresh at every floor, from the numbers on the dials.
The columns now descend through the Mill — floor after identical floor of linkage-and-gearwork, each refining the settings handed down by the last.
In truth — attention plus a small calculator makes one layer. Layers are stacked — dozens of them, close to a hundred in a big model. A word’s numbers are changed a little at every floor. Low floors seem to handle spelling and grammar; higher floors handle meaning and intent — or so the people who study these machines report. The adjustable parts are called parameters. A big model has hundreds of billions of them, and not one was set by hand. In training, the engine read a huge library and guessed each next word. Every time it guessed wrong, every screw was turned a hair. Do that a great many times and that is the whole of its schooling. A finishing school comes after — training on what people prefer — to make it helpful, honest, and polite.
This cabinet, for comparison, owns — counting-wheels, learned from a library of — verse lines.
At the mill’s outfall the engine renders its judgment — not a word, but a pressure upon every gauge in the house: a score for every slug in the catalogue at once.
In truth — at every step the machine gives a score to every token it knows — all 50,000-odd of them, most tiny. It does not pick the next word here. It rates them all. These gauges are real: they show this engine’s actual arithmetic for the words on the rack right now, and they answer to the boiler lever at Station VII if you move it.
For the first two words after a new card, the stop-mark and the full stops are held at zero, so a one-word card cannot be answered with silence. From the third word on, the engine may stop when it likes.
One slug must be drawn. The lots are cast in proportion to the gauges — and the boiler decides how strictly. Run cold, the favourite is all but certain; run feverish, and outsiders take their chance.
In truth — the next word is chosen by a weighted lottery — the trade word is sampling. Temperature changes the odds exactly as the drum shows. Turn it down and the favourite almost always wins: safe, but it repeats itself. Turn it up and the odds flatten: more surprising, and at the far end, nonsense. That is why the same question, asked twice, can get two different answers. Try the lever — the gauges change at once.
The chosen slug is stamped upon the tape — and carried straight back to the rack, where it joins your question. The whole engine then turns over again, from card to lottery, to choose the word after. One revolution, one word.
In truth — a language model writes autoregressively: one token per pass, each pass re-reading all that came before. There is no finished answer waiting inside; each word is minted the moment before you read it. And the engine stops as it speaks — by predicting a special stop token (printed here as ∎), whereupon it rests.
A real engine keeps notes (a cache) so it need not redo old sums on every turn. But in principle, every turn reads the whole rack again from the start.
Designed and built in a Victorian burst of enthusiasm for Shakespeare, the press sets the engine running line upon line — fourteen to the sheet, quatrains and couplet — in the hope, sincerely held and entirely vain, that new poetic masterpieces might emerge.
In truth — nothing new can come out. Every line is a walk through the sonnets the machine has read, joining them at words Shakespeare happened to use twice.
The engine read the poems line by line, as they were printed in 1609, so it has learned roughly how long a line runs and where one ends — it even has a mark of its own for the ending, ⏎. What makes the whole sheet look like a sonnet is not the engine but the harness: the row of cams above it, which supplies the fourteen lines and the rhyme scheme.
Each cam is one rule laid on the lottery. It either strikes words out of the drum before the draw, or tilts the odds toward the ones it wants. (The trade calls this constrained decoding.) A cam can forbid and it can steer — it cannot invent. Only the counting-wheels ever propose a word. If the engine has never read a rhyme for compass, no rule can produce one, and the margin will say so with a ✗. Watch for that failure: it is the whole point of the exhibit.
The trade calls this whole panel the harness, and every modern engine has one. When a chat machine refuses a question, or answers in a tidy list, or knows to stop talking — that is mostly harness, not intelligence.
Two things are worth taking away. First: the harness can shape what the engine says, but it cannot add anything the engine does not know. Second: what the harness can achieve depends entirely on how much the engine has read. Throw the Rhyme Cam at fifteen sonnets, then at all 154, and count the ✗ marks for yourself.
Rhyme is judged as Shakespeare judged it — by sound, allowing his own old pronunciations (love with prove, eyes with lies), which are marked in the margin as such. A word never rhymes with itself. And like him, the press does not stop for an imperfect rhyme.
The press runs on the same boiler as the crank — the lever at Station VII — so your sheet is printed at whatever setting stands there.
When the sheet is full, the concordance beneath it underlines every run of three or more words that appears, in that order, somewhere in the sonnets, and names the line it came from. It then tallies how much of the sheet is quotation and how much is the machine's own stitching. That is the question every reader of any engine's output ought to ask — and here it can be answered exactly, because everything the machine has ever read is on this page.