Language models need new ways to scale as more parameters, more data, more compute hit physical limits. We're taking a different approach: weights constrained to rotations rather than free matrices, so a layer can't blow a signal up or collapse it; and tokens stored as phases, compressing the embedding to about half the numbers a full vector would take.
The core idea: force each projection into a cascade of Givens rotations wired in an FFT butterfly pattern. A standard weight matrix can be anything — and much of that freedom goes to waste, spent rediscovering structure that every layer needs anyway. GRT builds that structure in from the start — fixed butterfly wiring, rotations instead of free matrices — and learns the angles, scales, and coupling on top.
Concretely, GRT reads the signal two numbers at a time as a polar pair — a magnitude r and a phase φ, written (r, φ), equivalently the Cartesian (x₀, x₁) = (r·cos φ, r·sin φ). That pair is the live activation the cascade transforms, not something stored.
The embedding stores a token the same way — as phases, not 256 loose numbers. With d = 256 that's 128 pairs, and it keeps one angle θ per pair plus a single magnitude r shared across the whole token. A stored pair is just (θ, r), and its Cartesian form x₀ = r·cos θ, x₁ = r·sin θ is rebuilt every forward pass, never saved.
The network keeps something different — not points but rotations (those come next). When the model runs, the embedding seeds the opening pair — stage 0's input — and the cascade turns and stretches it from there. The diagram below lays out what it stores; the matching activation half — those weights in motion — comes at the end of the section.
What a standard layer keeps (four dense 256×256 matrices + a [50000, 256] embedding) vs what GRT keeps (small angle / scale / coupling tensors), sliced down to one stored weight. The activation half — that weight transforming a token coord — is at the end of this section.
A projection (Q, K, V, O) uses the very same atom, but a stored weight here carries more than the embedding's lone angle. A standard layer stores a dense 256×256 grid of numbers; GRT stores, per pair and per stage, a phase θ, a scale pair (s₀, s₁) — held in the scales tensor and kept positive by softplus — and a small coupling W. The difference from the embedding: there (θ, r) is a point that becomes the activation, whereas these are the weights that turn the activation pair, stretch it along each axis by the scales, and let the coupling nudge the angle by the pair's own geometry.
One slot's stored weight is three learned things: the phase θ, the scale pair (s₀, s₁) (in the scales tensor, kept positive via softplus), and the coupling W. All are fixed during a forward pass and change only through training — unlike the embedding's per-token (θ, r), which is looked up by token id.
To make this structured representation dynamic, we introduce a Clifford-feature coupling. This coupling reads the length, direction, and signed area of each coordinate pair directly from the input signal to modulate the rotation angles on the fly.
A 2-D rotation is a phase shift; the cascade wires those rotations in an FFT butterfly.
That single idea buys structural guarantees. A rotation can only turn a signal, never amplify or collapse it, so the rotation backbone is length-preserving by construction — and that holds at initialization, before a single step of training. A standard layer has to learn to stay well-behaved, leaning on careful initialization and normalization to keep signals from blowing up or fading out; here that stability is built into the shape of the weights. Changing a signal's magnitude is left to the scales (s₀, s₁) alone — an explicit, learned knob — rather than something that can creep in through every weight. The model can still amplify or shrink, but only where it deliberately chooses to.
During training, gradient descent adjusts the rotation angle for each pair at each stage, and the scale parameter for each coordinate. The butterfly routing is fixed wiring with zero learned parameters. The Clifford coupling weights are learned. They determine how each pair's geometry affects its neighbors' rotations. The model's knowledge lives in angles and scales, not in dense grids of numbers.
Now the activation half. A token coord (x₀, x₁) — rebuilt from the embedding — runs through one stage: the coupling reads four features and nudges the angle (θ_used = θ + Δθ), the rotation turns it, the scale pair stretches it, and the butterfly re-groups for the next of 11 stages.
Reading each pair as a magnitude and a phase buys two things a plain vector doesn't. The two halves carry different kinds of information: the magnitude says how much a feature is present, while the phase says which configuration it's in, so the model can turn one up without disturbing the other.
And because each number is really a point on a plane, pairs interfere when they are added together: two pointing the same way reinforce each other, and two pointing opposite ways cancel out — constructive and destructive interference, exactly the way waves combine.
The heaviest parts of a transformer get lighter here: the embedding stores phases instead of a full vector per token, and each feed-forward network becomes a rotation cascade rather than dense matrices. That compactness is structural, not trimmed off afterward — how much it adds up to depends on the model's width and stage counts. The content-adaptive coupling (next card) spends parameters of its own to buy that adaptivity.
Every pair carries its own geometry: its energy, its direction, its twist. Those features feed back to steer the rotations, so the same coordinates are treated differently depending on what they actually carry.
A rotation can only turn a signal — it can never amplify it into noise or shrink it to nothing. Magnitude lives in a separate, explicit knob, not hidden inside a matrix. Training stays stable from the first step.
A standard transformer block has three learned projections: Q (queries, what to look for), K (keys, what to match against), and V (values, what to retrieve). These are matrix multiplies: each takes the input vector and produces a new vector. Attention then compares queries to keys, weights the values, and an output projection O projects the result back. After attention comes a feedforward network (FFN) that expands and contracts the signal.
Each projection runs across multiple heads. The model splits the d=256 vector into 4 heads of d=64 each. Every head has its own independent Q, K, V, and O parameters. Head 0 might learn one pattern while head 1 learns something completely different. They run in parallel and concatenate back. The mixer then crosses information between heads.
But strip away the names and every one of those projections is the same operation: a dense matrix multiply, \(y = Wx\). \(W\) is a grid of \(D\times D\) numbers, and each entry is a single scalar weight — a magnitude knob that answers exactly one question: how much? A weight can make a feature larger or smaller, but it has no native sense of direction. It cannot turn a signal, only resize it. The model's entire vocabulary is magnitude: \(D^{2}\) independent multipliers, each free to become anything.
That freedom is also the liability. Those matrices live in \(GL(N)\) — the set of all invertible grids, with no structural constraint on what they can do to a signal's size. Nothing stops one layer from amplifying it a hundredfold, or shrinking it toward zero. Stack twelve such layers and you meet the two failures every transformer trainer knows: exploding gradients, where magnitude runs away and the signal blows up, and vanishing gradients, where repeated shrinking flattens the signal until nothing can learn. Residual connections, layer normalization, and careful initialization all exist to fight this — but they are patches on a primitive that is, by design, free to explode or collapse.
GRT keeps every part of that architecture — the same Q, K, V, O, attention, FFN, and residual stack — and changes only the primitive. Where a standard layer multiplies by a dense matrix, GRT applies the stored rotate-and-scale from the storage diagram above: each pair is turned by a learned angle, then its two coordinates are stretched by a learned scale each. So what is that stored weight, exactly?
The phase-shift a pair gets isn't fixed. The learned weight contributes a base phase θ (stored), and the pair's own geometry adds a small nudge. This is the one place the polar story reaches back into Cartesian: the nudge is computed from the pair's raw components x₀, x₁, not from (r, φ) directly. GRT reads four Clifford features off the pair — energy x₀²+x₁² = r², the two direction components x₀ and x₁, and the signed area it forms with its neighbour x₀x₁′ − x₁x₀′ = r·r′·sin Δφ — multiplies them by learned coupling weights, squashes with tanh, and adds the result to θ. So the effective phase-shift is θ = base phase (stored) + nudge (from the pair). Rotation, scale, and the embedding all live cleanly in polar; the Clifford coupling is the exception that reads the Cartesian x₀, x₁.
Every pair's nudge comes from one matrix, the coupling W, shaped [128, 32]. The 128 rows are feature-major — the four Clifford features stacked, each spanning all 32 pairs: f₀ for pairs 0–31, f₁ for 32–63, f₂ for 64–95, f₃ for 96–127. The 32 columns are the output slots.
Slot p's nudge reads the full column W[:,p] (128×1). The four highlighted cells are the self 1×4 — rows {p, 32+p, 64+p, 96+p}.
To produce slot p's nudge you read its whole column, W[:, p] — a 128×1 vector, so every pair's features can influence pair p. The four entries from pair p's own features sit at the strided rows {p, 32+p, 64+p, 96+p} — one per band — the “self” 1×4. Nothing extracts that slice in practice: the forward pass runs the whole product Δθ = tanh(F·W), all 128×32 at once. The 1×4 is just where a single pair's own contribution lives.
The neighbour never appears as a separate weight. Pair p's f₃ is built before W by a cyclic roll — f₃[p] = x₀[p]·x₁[p+1] − x₁[p]·x₀[p+1] — so the cross-pair (bivector) interaction costs zero parameters; only its weighting by W is learned.
Each learned per-pair transform rotates by an angle \(\theta\), then scales the two coordinates by \(s_0\) and \(s_1\). When those two scales are equal it's a uniform scale-and-rotate — an element of \(SO(2)\times\mathbb{R}_{+}\), isomorphic to \(\mathbb{C}^{*}\) (the nonzero complex numbers under multiplication, via \((s,\theta)\leftrightarrow s\,e^{i\theta}\)). But GRT's scales are generally anisotropic (\(s_0\neq s_1\)), so the transform is really a general orientation-preserving linear map — an element of \(GL^{+}(2,\mathbb{R})\) — with the tidy complex-number case as a special case. Either way GRT only ever multiplies these per-pair maps; it never adds two complex pairs, so any "addition-like" behavior comes from the real residual stream and normalization, not from a complex field.
And the imaginary unit \(i\) is never actually used: there is no complex datatype, no complex autograd, and no conjugation. Each per-pair map expands term-by-term into four real entries — a rotation \(\left(\begin{smallmatrix}\cos\theta & -\sin\theta\\ \sin\theta & \cos\theta\end{smallmatrix}\right)\) followed by a per-axis scale \(\mathrm{diag}(s_0, s_1)\) — exactly the Givens rotation and scale the code already applies. The complex picture describes the isotropic case faithfully, but it's a lens, never an ingredient.
That was one stage. The full Q projection chains 11 of them together — each with its own learned angles, scales, and coupling, and a butterfly re-pairing in between that regroups the coordinates (a different stride each stage).
One polar pair plus a magnitude looks tiny. The cascade is what compounds it. Each of the 11 stages turns the pair by a learned angle and stretches its two coordinates, then the butterfly routing hands it to a new partner, so a local pair reaches across the whole 256-d vector within a few stages. The energy coupling ties pairs together: each pair's energy r² and the signed area it shares with its neighbour feed the phase nudges, so one pair's geometry bends another's. With no normalization between the stages, the phases pile up and the magnitudes spread freely, turning the embedding's uniform ring into a structured, content-shaped representation. The primitive stays one polar pair plus a magnitude — the structure is what the cascade builds out of it.
The uniform ring those plots start from is the embedding's doing, not PearlNorm's: every pair of a token shares one magnitude, so block 0 enters as a near-perfect circle (radius ≈ 0.79). PearlNorm runs before each sublayer, but it doesn't re-impose that ring — it rescales by a single per-token factor read from the pair radii so depth can't let magnitude run away, and its learned per-coordinate gain leaves the radii unequal. The cascade is free to spread the radii apart, and that spread is the information.
Q projection runs 11 stages.
Each stage shuffles pairs, reads Clifford features, nudges angles, rotates, and scales. The plot's Z=0 is the embedding's ring (radius ≈ 0.79, pre-PearlNorm); PearlNorm re-bases it before the first stage, and the learned parameters then sculpt it into structure.
The output is a set of query vectors, one per head. Queries encode what each token is looking for. You just saw one pair traced through all 11 stages in the concrete example above. The plot shows all 128 pairs doing the same thing simultaneously.
A note on the axes: every plot here is drawn in the reconstructed Cartesian coordinates x₀, x₁ — because that's what you can plot. But those are not what's stored. What's stored is the polar pair (θ, r) — an angle and a magnitude — and the x₀, x₁ you see are rebuilt from it each step by the bridge: x₀ = r·cos θ and x₁ = r·sin θ. The curves move because the stored angle and magnitude move; the Cartesian trace is just that bridge made visible.
Four heads, one shared wiring.
The d=256 vector is split into 4 heads of d=64. Every head runs the same fixed butterfly wiring and the same 11-stage schedule — the routing is identical and carries zero learned parameters. What differs is everything the model actually learns: each head has its own angles, Clifford coupling, and scales for Q, K, V, and O. Same plumbing, four different learned behaviors — so each head ends up attending differently.
Each head's Q projection (11 stages) for 'earth', drawn alone then combined. Identical butterfly wiring; different learned angles, coupling, and scales per head.
Interactive: 4 pairs per head lit in their head color, the other 112 faded back. Drag to rotate and watch each chosen pair twist through its 11 Q stages.
K runs 11 stages.
Same cascade structure, different learned angles and scales. K produces keys, which determine what each token attends to. K learns different parameters than Q because keys and queries serve different roles in attention.
The key cascade is structurally identical to Q: 11 stages, butterfly routing, Clifford coupling, per-pair scales. But the learned angles are completely different. The model discovered through training that keys need to point in different directions than queries for attention to work. Two cascades with the same wiring, solving different problems.
V runs 11 stages.
Same cascade again, different parameters. V produces values, which carry the content that attention will retrieve. V learns yet another set of angles and scales because values carry meaning, not questions or labels.
After V, the model has three representations of the input: queries (what to look for), keys (what to match against), and values (what to pull out). Attention combines them.
Softmax attention weights V.
This is standard transformer math. \(\mathrm{softmax}(QK^{\top}/\sqrt{d})\) determines how much each token attends to every other token. The values get weighted-summed by these scores. No rotation stages here, no butterfly, no Clifford coupling. Just dot products and softmax.
If Q asks the question, K holds the label that answers it. The dot product between a query and a key tells the model how relevant two tokens are to each other. Softmax turns those scores into weights that sum to one. Those weights multiply V, the values, to produce the attention output. V carries the content that gets pulled out.
Nothing about attention changes in GRT. Three rotation cascades produced Q, K, and V. The matching and retrieval is the same as any transformer.
Q, K, V cascades (all 4 heads) meeting at the attention layer for 'earth'. Marker color = cos of the Q–K angle: blue = aligned (high attention), red = opposed (suppressed). This is the geometry feeding softmax(QKᵀ/√d)·V — the same attention equation as any transformer, fed by rotation cascades instead of weight matrices.
Per head, the same picture: each head's Q, K, and V cascades meeting at its own attention layer. The marker color is the Q–K alignment that drives attention — and it differs head to head, because each head learned different angles.
Q/K/V→attention for each head separately, then combined. Blue = aligned (high attention), red = opposed.
O projection runs 11 stages.
The attention output gets projected back to d=256 through another rotation cascade. This is the fourth and final cascade in the attention sublayer.
This is the sharpest break from a standard transformer. A normal block sends the mixed attention values back to the residual stream through a single dense projection — one flat \(W_O\) matrix, \(O(D^2)\) parameters, applied once. GRT replaces that one map with the same structured cascade it uses for Q, K, and V: the freshly mixed token states are pushed through 11 stages of butterfly-partitioned Givens rotations and Clifford-coupled angle nudges, each followed by a learned scale.
Call it a multi-stage structured output cascade. A standard layer's output map is a single flat matrix; to deepen post-attention processing it has to stack more of them. GRT instead refines, rotates, and re-projects the mixed attention states through a multi-stage group-theoretic pipeline — deeper, more structured output processing than one dense map gives, built from stability-preserving rotations.
After O, the attention result is added to the input via a residual connection: x + attention(x). This is where the projection-level det > 0 guarantee can break at the Jacobian level, because the residual Jacobian I + J_f can go negative.
The same per-head view for the output cascade: each head runs its own 11-stage O projection on the attention result — identical butterfly wiring, different learned angles, coupling, and scales.
Each head's O projection (11 stages) for 'earth', drawn alone then combined.
That is the whole attention sublayer. Every plot so far showed all 128 pairs at once — organized chaos. To recap, here it is distilled: 16 pairs, 4 from each head, carried through Q, K, V, attention, and O on a single plot in their head colors, with the other 112 faded back. Same rotate-and-scale primitive, every stage, every head.
16 pairs (4 per head): Q/K/V → attention → O on one projection, head-colored, the rest faded.
Residual +, then PearlNorm.
The attention output is added to the input. PearlNorm runs again before the FFN — rescaling by a per-token factor read from the pair radii, not flattening the radii to equal. This is the last normalization before the feedforward network.
The signal is now a mix of the original input and what attention found. PearlNorm re-bases its overall scale to a stable starting point for the FFN — angle and the learned radius profile preserved, not reset to equal.
Mixer runs 1 stage across all heads.
Q, K, V, and O each operated within a single head (d_head=64, 32 pairs). The four heads ran in parallel but never talked to each other. The mixer fixes that. It runs one full rotation stage at d_model=256 (128 pairs) — the same content-adaptive primitive as a projection (Clifford coupling, Givens rotation, scale), now spanning all four heads, so it mixes information across head boundaries.
After the mixer, what head 0 learned is visible to heads 1, 2, and 3. The FFN gets the full picture, not four isolated slices.
The attention sublayer in full: Q/K/V→attention→O, then O is added back to the residual stream (top layer, real values = block input + O output). The mixer then runs its one rotation stage across all 128 pairs and is added again — crossing information between the four heads before the FFN.
FFN expand runs 21 stages.
The feedforward network is where the model does most of its thinking. Expand grows the signal from d=256 to d=1024. That means 128 pairs become 512 pairs. But PearlNorm ran at d=256, not d=1024. So the 256-dimensional signal gets padded with 768 zeros to reach 1024. Only 128 pairs carry signal at the start. The other 384 pairs are empty.
Through 21 rotation stages, the butterfly shuffles coordinates so the 128 active pairs mix into the 384 empty ones. The model fills the new space with learned transformations of the input. In a standard transformer this step alone is a dense matrix of 256 x 1024 = 262,144 parameters. Here it is 21 rotation stages.
RadialGELU gates the signal.
Between expand and contract, RadialGELU applies GELU to each pair's radius while preserving its angle. This decides which pairs carry signal forward and which get suppressed. It is the only nonlinearity in the FFN.
Some pairs get amplified, others damped, but no pair's direction is distorted. It is what lets the model decide which features matter and which to silence.
FFN expand → RadialGELU. The 512 pairs run 21 rotation stages, then RadialGELU gates each pair's radius (angle preserved) — the only nonlinearity in the FFN.
FFN contract runs 21 stages.
The 512-pair signal shrinks back to 128 pairs (d=1024 to d=256). The model compresses what it learned back into the residual stream's dimension. Another 21 rotation stages, different learned angles.
What the model discovered during expand gets distilled into the output. The same butterfly routing carries information from all 512 pairs back down to 128.
Residual +. That is ONE block.
The FFN output is added back to the residual stream. The block is done. Q (11) + K (11) + V (11) + O (11) + Mixer (1) + FFN expand (21) + FFN contract (21) = 87 rotation stages per block. Its output feeds into the next block, which repeats everything with its own learned parameters. There are 12 blocks. That is 1,044 rotation stages total — no dense weight-projection grids, only the small per-stage coupling matmuls.
Every rotation stage from block 0 stacked on one plot. Z=0 is the embedding's own uniform ring (one shared magnitude per token), not a circle imposed by PearlNorm. Then Q (11), K (11), V (11), O (11), FFN expand (21), FFN contract (21). Attention and mixer stages run over 128 pairs (4 heads × 32 at d_head=64); the FFN stages run over 512 pairs (d_ff=1024). The signal flows through the entire block in one continuous view.
Twelve blocks, each one full pass.
One block does three things — attention mixes information across tokens, the mixer crosses it between heads, and the FFN transforms it per token. Stack twelve and the model runs that cycle twelve times. Each block reads the stream, transforms it, writes back; the next block works from what the previous one left.
What stays identical across all twelve: the wiring. The butterfly schedule, the PearlNorm rhythm, the order of the sublayers — set by the architecture, not learned. What's different in every block: the angles, the scales, the coupling weights, and the residual gates. Each block learns its own parameters from scratch. The shape is shared; the contents are not.
Zoom into any one of those GRT boxes — Q, K, V, O, or the FFN's expand and contract — and it is not a matrix at all. It unrolls into the cascade we traced earlier: eleven rotation stages, each one a full activation, with a butterfly re-pairing in between.
Each GRT box = 11 rotation stages. One stage is one activation; the butterfly re-pairs the coordinates between stages.
That is the whole mechanism in one frame. The left half is the per-stage activation: read four Clifford features off the pair, normalize them, multiply by the coupling W to get the angle nudge Δθ, add it to the base angle, then rotate and scale. The right half chains eleven of them, the butterfly shuffling each pair into new neighbours so a local pair reaches across the head within a few stages. K, V and O unroll the same way; only the learned angles, scales, and couplings differ.
A block is the structure we just walked through, configured. The same shape repeats twelve times; what changes is the learned parameters inside it. Every number below comes from the config — a choice, not a constant.
Normalize, process, gate, add.
Every sublayer — attention, mixer, FFN — follows the same four-step cycle. PearlNorm normalizes a copy of the stream, so the sublayer sees a clean input no matter how deep we are. The sublayer does its work: rotates pairs, reads coupling features, scales. A learned gate decides how strongly to write the result back. Then the residual add accumulates it into the stream.
This cycle is why a twelve-deep stack is trainable at all. The rotations alone are well-conditioned — they can't amplify or collapse a signal — but that isn't enough across twelve blocks of accumulated writes. Normalize-a-copy gives each sublayer a stable input. The gate stops any one sublayer from dominating. The residual add keeps the history. Orthogonality is the floor; normalize-process-gate-add is the structure built on it.
The cycle runs three times per block — attention, mixer, FFN — and the block repeats twelve times. That's thirty-six rounds of reset, transform, gate, accumulate from input to output.
The 256-dimensional signal splits into four groups of 64 coordinates. Each group is a head. Every head runs its own Q, K, V, and O rotation cascades with its own learned angles. The four heads share wiring but nothing else — four parallel processors looking at four different slices of the pair stream.
The attention score between two tokens sums, over all pairs, how aligned that token's Q pair is with the other's K pair — each pair contributing the cosine of its Q–K angle gap — and softmax turns those scores into weights. So a head's "strategy" is which pairs it aligns for which tokens: align the angles on a pair and two tokens attend strongly there; push them ninety degrees apart and they don't.
In a standard transformer, a head's Q projection is a fixed matrix — the same map for every input. In GRT each rotation angle is nudged by what the pair carries (its energy, direction, and twist), read through the Clifford coupling. One head can run different strategies on different inputs. That's "content-adaptive" in practice — and the mixer adapts the same way, through its own coupling. The FFN's rotation angles are fixed; it adapts only through RadialGELU, which gates each pair by its radius.
The heatmap below shows real attention weights on "the earth is blue". Each head learns a different role: head 2 acts as a self-attention filter, suppressing cross-token mixing. Head 1 spreads evenly across all tokens. These are learned specializations — the architecture provides the mechanism, training decides who does what.
Every pair's Q and K end up at different angles because Q and K use independent learned rotations. Below: the angle gap for all 128 pairs across all 12 blocks. Blue means aligned (high attention). Red means opposite (suppressed).
Between blocks, information moves through a 256-dimensional residual stream — 128 pairs of angle and radius. Every block reads from it and writes back. Because each write is a rotation plus a scale, a single sublayer can only turn pairs and resize them. It cannot, in one step, collapse the signal to nothing or blow it up to noise. That's a stronger guarantee than a dense matrix gives.
PearlNorm runs on a copy. The sublayer sees a scale-normalized input — every coordinate divided by one per-token factor read from the pair radii — so depth doesn't drift the input distribution. But the raw, unnormalized stream passes straight through the residual connection. Local normalization for stability; global accumulation for memory. The stream keeps everything; each sublayer just borrows a normalized view.
Each sublayer's write is scaled by a learned gate before it's added. In block zero, attention writes at about 0.99 — nearly all of itself, because attention refines what's already there. The mixer at about 0.97. The FFN at about 0.79 — it's the largest transform in the block, so the model learns to gate it down to keep the stack stable. The gates drift across the twelve blocks as the model rebalances who contributes.
This is the part only depth produces. Block zero's stream is almost the embedding — phases and magnitudes. Block eleven's stream is the model's final representation, shaped by thirty-six rounds of normalize-process-gate-add. Across the stack, pair radii differentiate: some pairs grow, some shrink, as the model decides which features matter. Pair angles settle into stable configurations. The stream doesn't get overwritten; it gets sculpted. By the end, the same 256 dimensions carry something the embedding never could.
A single sublayer can only rotate and rescale — it cannot catastrophically distort what's already there in one step. Over twelve blocks the stream is reshaped, not erased.
GRT is a transformer with one thing changed: every dense weight matrix is replaced by a cascade of rotations. The routing, the attention, the residuals, the depth — all standard. The primitive is different. A rotation is one angle; it can turn a signal but never break it. That single swap is what makes the model space-efficient by math, content-adaptive where it needs to be, and stable by construction.
Benchmark results at scale — HellaSwag, MMLU, ARC, PIQA, perplexity.