How AI Emerges, Through the Lens of Emergence

An ordinary city street generated for the E8 Lens: shops, pedestrians and traffic

Somewhere between a pile of numbers and a machine that can tell you what a street vendor is selling, in Hindi, something appears that nobody built. No line of code says “this is a market”.

Let’s explore how AI emerges, through the same lens this site uses for water, cells and cities: what had to exist first, what the parts allow, and what the whole does that no part does.

The short version: an AI like ChatGPT, Gemini or Claude does one thing, predicting the next word one piece at a time, and nearly everything else it seems to do grows out of getting very good at that.

We will stay on one ordinary street the whole way, drawn for your country this week (from your timezone, never your location), and look at it from every side: through a human mind, as the data it gives off, as the material machines learn from, at the moment ability appears, as a question put to a model, and finally as the shape it takes inside one. First, what emergence means here.

The two-minute version

A modern AI model is a very large set of numbers, tuned on huge amounts of human text and images, that predicts what comes next, one piece at a time.

  • It is autocomplete, scaled up. Your phone suggests the next word from a handful of numbers; a large model does the same with billions of them, after reading a large share of everything people have put online.
  • People made its lessons. Boxes drawn around cars, CAPTCHA squares ticked, chatbot answers rated better or worse: small, repetitive work, much of it done by people who never knew they were teaching a machine.
  • Training is one small move, repeated. Guess the next piece, measure how wrong it was, nudge the numbers. Trillions of times.
  • Nobody writes the abilities in. Translation, arithmetic, a sense of what a market is: they emerge from scale. Smooth underneath, sudden on the surface.
  • What it keeps is a map, not a list. Ideas end up near related ideas, across languages: “market”, बाज़ार and 市场 land in one place.
  • Through the lens of emergence, AI sits on every floor below it: chips (E0–E3), people’s labor (E8), languages (E9), companies (E11), laws (E12), and an internet’s worth of shared memory (E13).

Below, each of these happens on one street you can try for yourself.

Emergence, in brief

Emergence is the name for a specific gap. Hydrogen and oxygen are fully described by their own rules, and nothing in those rules tells you that combining them produces something that expands when it freezes. The lower level permits water; it doesn’t predict it.

That gap — between what the parts allow and what the whole actually does — shows up everywhere: a cell doing what no molecule does, a market setting a price no trader chose, a language no speaker designed.

Emergence is the pillar that works through this definition, and Abstraction explains the other half of it — that emergence is only visible through a chosen lens, and every lens shows one aspect while hiding the rest.

The two pillars together are the operating instructions for everything else on this site: pick a thing, ask what had to already exist for it to appear, and notice which lens you used to see it.

The E-levels are the index that makes those questions comparable across wildly different subjects.

The Arena lays out fifteen of them, E0 to E14, each defined by the transition that had to happen before the next could form: fields → particles → nuclei → chemistry → information-carrying molecules → cells → organisms → ecosystems → minds → cultures → institutions → organizations → states → civilizations → whatever planetary-scale coordination turns out to be.

A level is both a set of rules (the arena) and a player in the level above it. The numbers mark dependence, not rank; you are made of all fifteen simultaneously, and the Arena’s point is that you only ever see one at a time.

From there the paths split: Parameters covers why each level’s past choices constrain what the next can become, The Sensing Surface covers why the mind doing the looking is itself shaped by what it looks at, and the E14 Explorer lets you move the lens yourself.

One street, through one mind

You can climb that ladder yourself: the Arena, floor by floor — each of the fifteen levels with its rules, the things that exist on it, the dead ends that never leave it, and where it turns up in your own week. Then come back down to E8, and look through one mind at a time.

E8 Lens. The ladder climbs from E0 to E8, then one street is read through a mind. Change who is looking, slide the ladder, or ask what that mind is not seeing. A model of what may become salient to an observer, not what anyone objectively sees.

Hold on to this street; every section below comes back to it. It is not a photograph. One model drew it, another mapped every thing in it and placed each on the ladder, and a third imagined how each of those minds would read it. That is the question underneath this piece: how did machines come to do any of that? Start where they start, with data.

What the street says

The same street also talks. Every exchange in it leaves a trace: a price asked and settled at a stall, a card tap that leaves for a bank, two neighbors trading news about next week's road closures, a traffic light broadcasting one rule to everyone at once, bread announcing itself to anyone downwind.

Each exchange lives on a floor of the ladder. The smell is chemistry, the eye contact at a crossing is minds, the haggling is organizations, the license plate is the state.

Most of it does not stay on the street. Below is the data this street generates, drawn over the same picture and the same things the lens picked out, and followed to where it goes once it leaves the frame.

Street Data. The same street as the E8 Lens, read for the data it generates. Tap a line or a highlighted thing; the ladder sorts every exchange by the E-level it lives at.

How a machine learns

Every model behind these pictures learned to see from examples. Someone drew a box around a taxi and typed "taxi". Someone ticked the squares with buses to prove they were human. Someone read two chatbot answers and picked the better one. And billions of pictures and sentences had pieces hidden for a model to guess back.

None of that is abstract: it is small, repetitive work, much of it done by people who never knew they were doing it. Try each job on the same street, then see what this street's own data goes on to train, and what the models behind these pages got wrong because of what they were trained on.

How a machine learns this street. Pick the squares, tag the picture, rate a chatbot, fill the gap. Nothing you do here is sent anywhere.

The ladder of learning

Those are the jobs you can see. The rest of AI's training runs up and down the whole ladder.

At the bottom, image generators learn by undoing noise, and controllers for fusion reactors practice in simulators before they touch real plasma. In the middle, perfumers' words teach models to smell, fifty years of lab work taught one to fold proteins, and a walker learns to cross a road from nothing but points.

Higher up, written rules shape what a chatbot will say, every tap on a feed is a label, the law decides what may be learned from at all, and models now learn from text that other models wrote.

Each floor has its own way of teaching a machine. Here they are one rung at a time, all on the same street.

The ladder of learning. Pick a floor, E0 to E14; each shows one way machines are trained, on the same street.

When it emerged

None of that training contains an ability. No one wrote "this is how to translate" or "this is what a market is". The machine only ever does one small thing: guess the next piece, measure how wrong it was, and nudge its numbers.

It is your phone’s autocomplete, the word it suggests as you type, and the tiny model below is exactly that, guessing one letter at a time.

Repeated a few thousand times on one street's text, that gives spaces, then common words, then the street's own words, and then it runs out of street. Repeated trillions of times on a large share of everything people have put online, the same nudges add up to translation, arithmetic, and a sense of what a market is in many languages.

The scale is hard to hold in your head. The largest of Meta’s Llama 3.1 models, among the biggest whose training is public, took about 31 million hours of data-center chip time: a single chip that started before Tutankhamun was born would only now be finishing.

Underneath, the change is smooth and steady. On the surface, abilities seem to switch on. That is emergence, the same move the ladder makes from molecules to minds. Watch a tiny one happen here.

When it emerged. A tiny model trains live in your browser on this street's text; then the scale it took for the big ones.

Asking the model

Then someone asks it a question. "What is that man selling?" "Is it safe to cross?" "Who is she?"

The model has never seen this street. It answers from everything above at once: boxes that strangers drew, squares that people ticked to prove they were human, billions of sentences it learned to finish, answers that raters preferred, rules someone wrote down, and data the law kept out.

Every reply is the ladder of learning, used in a second. Ask it something, and watch which floors it climbs to answer.

Asking the model. Pick a question; each step shows what it looks at and the kind of training that step depends on.

What it understood

So what does a model actually hold, after all that? Not the street. A space.

Every word, picture and idea it has learned becomes a point in a space with thousands of dimensions, and understanding is where things land in it.

A word-vector model from 2014 placed "market" next to stocks and investors and "street" next to Wall Street, because that is what the English news it read was about; it had no point at all for बाज़ार or 市场.

A frontier model today places this street's food, bank, traffic rules, trust and the monetary system on one connected web from chemistry to civilization, and the same idea in Hindi, Chinese, Arabic and English lands in nearly the same place. Nobody wrote that in. It emerged.

What it understood. Positions are each model's own vectors, projected to three dimensions. Drag to turn the space; switch between 2014 and 2026.

Back to the lens

Run the site’s three instructions on AI itself. Pick a thing: a model that can read a street it has never seen.

Ask what had to exist for it to appear: physics tamed into chips that flip billions of switches a second (E0–E3), people drawing boxes and rating answers (E8), languages to learn from (E9), rules and laws deciding what it may learn (E10, E12), companies and paid labor to build and run it (E11), and an internet’s worth of text, the nearest thing we have to a civilization’s memory (E13).

Then notice the lens. Everything here was seen through one street, and a street hides as much as it shows: the energy bills, the data centers, the workers who never appear in the picture.

AI did not arrive as a single invention. It emerged, like water and markets and minds, from layers that permit it without predicting it. Whether it is a step toward whatever E14 turns out to be is the open question the ladder ends on.

If you have followed this far, an idea of how AI emerges has just emerged in your mind. It has already connected with everything else you knew. That is what AI labs are trying to make happen inside their models. Now you know.


A note on the images and data. The street, its data and the examples in this explainer are personalized to make them relatable to you: your country is picked from your timezone (never your location), along with the season and that week’s positive news. They are regenerated every week, so the next time you visit, the street may look different and some of the details may be new.