Clicked Gallery

What is a Looped Transformer?

Highlighted from a real engineering doc. Explained by Clicked.

Used in a sentence

Engineering Notes · AI Systems

The new model is reported to be a looped transformer, running two passes through its layers instead of three.

The reader highlighted one word in the docs. Clicked explained the technical term β€œlooped transformer” in simple terms:

Explained in three depths

Same facts, different vibe β€” Slang mode 😎

The Clicked way

●○○

Overview

A looped transformer is an AI model that sends each question through the same block of layers more than once, instead of giving every step its own layer. A layer is one step of computation: data goes in, the layer's parameters transform it, and new data comes out. The parameters are numbers learned in training. An ordinary model with 24 layers runs the data through each layer once. A looped model with 12 layers runs the data through all 12, then sends the output back through the same 12. The parameters do not change between passes; the data does, because the second pass starts from the first pass's output. The same loop runs in training, where the parameters are being adjusted, and in inference, where the parameters are fixed. A looped model therefore gets 24 steps of processing from 12 layers' worth of parameters, which is half.
●○○

Overview

A looped transformer is AI that runs a question through its layers two or three times instead of having more of them. A layer is one round of number-crunching: data in, numbers learned in training crunch it, data out. The parameters are those numbers, and they're what eats memory. Say the model has 6 layers. A question runs through all 6, then heads straight back into the first layer for a second lap, and a third. The numbers stay put between laps; the data moves on, because lap two starts where lap one got to. Training runs the same laps while the numbers are still being nudged; inference runs them once the numbers are set. That's 18 rounds of crunching, 6 times 3, from 6 layers of parameters. The catch: three laps take three times as long. 😎

A quick take β€” often all you need.

●●○

Detail

A looped transformer is an AI model that sends each question through the same block of layers two or more times, each pass taking the previous pass's output as its input. A layer is one step of computation: data goes in, the layer's parameters transform it, and new data comes out. The parameters are numbers learned in training. Depth is how many such steps a question gets, and looping buys depth without new parameters. The same loop runs in training and in inference. In training the parameters are adjusted after each answer is checked, so only parameters that work when applied again to their own output survive. In inference the parameters are fixed, and only the data changes from pass to pass. Google Research measured what looping buys (ICLR 2025). On adding eight three-digit numbers, a 1-layer model got 0.1% right, the same 1-layer model trained to loop 12 times got 99.9%, and a 12-layer model got 100%. The cost is time: each pass is a full run through the block, so a 12-layer model looped twice takes about as long as a 24-layer model. The loss is room for facts. In the same study a 12-layer model looped twice beat a 24-layer model on short school-style maths questions written in words, 34.3 against 29.3. The looped model lost on trivia questions, 9.3 against 11.2, because facts are stored in parameters and the looped model had about half as many. Google fixed the number of loops before training. Some looped models are trained with the count varied, so the number of passes can be set at inference. Either way the loop runs inside the model, before the first word of the answer, which is what separates looping from chain-of-thought, where the model writes its steps out as words.
●●○

Detail

A looped transformer is AI that pushes a question through the same layers again rather than through new ones. A layer is one round of number-crunching: data in, numbers learned in training crunch it, data out. The parameters are those numbers, and they're the bit you pay for in memory. So stop adding layers and reuse the 6 you have, three times round. The numbers stay put between laps; the data moves on, so lap two picks up where lap one left off. Training and inference do the same three laps. In training the numbers get nudged after each answer is marked; in inference the numbers are set. 18 rounds of crunching, 6 times 3, for 6 layers of parameters. The bill still arrives: three laps is three lots of computing, about what an 18-layer model would take. The trade: in the Google study the looped models did better on sums and on answering from a passage, and worse on pub-quiz facts, because facts live in parameters and a looped model has fewer. Some looped models are taught with the lap count mixed up on purpose, so you can pick 2 laps or 3 at inference. Chain-of-thought is a model spelling its steps out in words before it answers. Looping is the extra steps happening internally, nothing spelled out, and the lap count is becoming the thing people ask about a new model, not only the parameter count. 😎

Want more? One click digs deeper.

●●●

Analogy

A looped transformer is an AI model that sends a question through the same block of layers more than once, so that the second pass works on what the first pass produced. Proofreading your own essay works the same way. You read the essay once and fix what you can see. Then you read it again. Your eyes and your knowledge have not changed, but the essay has, and the fixes from the first read expose mistakes that were hidden before. That second read is the loop: you are the block of layers, and each read is one pass. The alternative is to hand the essay to two different readers, one after the other. Each reader brings knowledge you lack, so two readers are the stacked model, with a fresh set of parameters at every step. Two readers catch more factual slips, and two readers cost two people instead of one. Where the picture breaks: a second read costs you only time, while a second pass through a model costs as much computing as an extra layer would.
●●●

Analogy

A looped transformer is AI that runs its data through the same layers again and again instead of adding more layers. A layer is one round of number-crunching, and the numbers it crunches with are the parameters, learned in training. A tumble dryer is the same deal. The clothes come out damp. Nobody buys a second dryer and moves the load across. You shut the door and run the same drum for another 20 minutes. Same machine, same settings; only the clothes are different. The drum is the block of layers, one cycle is one pass, and buying a bigger dryer is stacking more layers. Once the clothes are dry, another cycle does nothing and still costs electricity, which is the limit on how many loops are worth running. Where it comes apart: a dryer isn't smarter on cycle two, it just gets more time, while a looped model's second pass starts further along than its first. And a dryer never had any facts to forget. 😎

Unfamiliar concept? A real-world example makes it click β€” fresh analogies on tap.

AI explanations may contain errors · Not professional advice

Formal definition β€” The same term, explained the usual way

A looped transformer (also described as a recurrent-depth or weight-tied transformer) is a transformer architecture in which a block of one or more layers is applied repeatedly to its own output, sharing the same parameters across iterations, so that effective depth equals the block depth multiplied by the number of loops while the parameter count stays fixed. The idea descends from Universal Transformers (Dehghani et al., 2018). Compute per forward pass matches a non-looped model of the same effective depth (the iso-FLOP baseline), so looping trades parameters for serial computation rather than reducing inference cost. A common variant, middle looping, keeps separate prelude and coda layers and loops only the middle block. Saunshi et al. (ICLR 2025) report that looped models match iso-FLOP baselines on reasoning tasks with far fewer parameters, show worse perplexity and closed-book recall, and can simulate chain-of-thought with latent rather than written intermediate steps; the loop count can be varied at inference time.

Want Clicked to explain terms like “looped transformer” directly in your browser β€” including on PDFs?

Add to Chrome β€” Free

50 free Explanations · No credit card required