Research Gallery — Ouro × brain encoding

Persistent lab gallery. Every plot we make gets added here with a label so we can always come back. Goldstein et al. 2025 replication, then extended to the Ouro-2.6B looped transformer. Language ROIs: TP, IFG, aSTG. Auditory control: mSTG. Reading: in language ROIs the encoding-peak lag drifts LATER as model depth increases; in the control it should not. All 'ours' panels regenerate from frozen_poster_data/ — the isolated dataset the poster + research were built on.

1 · Paper — Goldstein et al. 2025, Fig 3 (GPT-2 XL)

The published result we are replicating. Color = layer 1→48 (red→blue). Paper r: TP +0.93 · IFG +0.85 · aSTG +0.92 · mSTG −0.24 (ns).

Fig 3 top row — per-layer encoding curves + peak-lag dots
As provided (photo of the paper). Language panels: stacked dots march rightward (later lag) from red→blue. mSTG control: dots stay pinned near lag 0 — no hierarchy. (Bottom-row scatter is in the PDF; can be added on request.)

2 · Ours — GPT-2 XL replication

Our pipeline (PCA-50 + OLS, predictive word-alignment). Should reproduce the paper.

All four ROIs on one axes — peak lag vs GPT-2 layer (the clean summary)
The whole replication in one panel: each ROI as colored points + a dashed regression. TP +0.88 · IFG +0.82 · aSTG +0.80 (language ROIs rise) · mSTG control −0.26 (ns) stays FLAT. The flat green control is the language-specific signature — exactly what the paper shows. Raw-argmax peak read-out (matched to the Ouro combined plot for a fair side-by-side; a hair off the parabolic papermatch +0.89 etc). Regenerated from best_run via analysis/fig_gpt2_combined_scatter.py.
Gold standard — papermatch (clean regression-line Fig 3)
TP +0.89 · IFG +0.83 · aSTG +0.81 · mSTG −0.26 (p=.07 ≈ paper's −0.24). The gold-standard presentation: clean single-row Fig-3 scatter with a bold dashed regression line + big bold r per panel (same data as best_run, cleaner than the raw fuzzy scatter). Posterior-STG proxy control + sub-bin parabolic peak interpolation. Regenerated from best_run via analysis/fig_papermatch.py.
Definitive replication — top: curves, bottom: peak-lag-vs-layer scatter (window 0–1000 ms)
IFG +0.82 · aSTG +0.80 · TP +0.88 (all significant) · mSTG control −0.58 (ns). Auditory control = re-gated Heschl electrodes. Matches the paper's direction and significance pattern.

3 · Ours — Ouro-2.6B (looped / recurrent transformer)

The novel extension. Ouro reuses the SAME weights across 4 recurrent passes; its "depth" is unrolled iteration: effective depth = pass×48 + layer (0→191). Color = effective recurrent depth.

All four ROIs on one axes — peak lag vs effective recurrent depth (the contrast)
Same one-panel summary as the GPT-2 plot, directly comparable. Language ROIs rise (IFG +0.66 · aSTG +0.60 · TP +0.25 ns) — BUT the green mSTG control is now the STEEPEST line of all (+0.84), not flat. That single difference vs GPT-2 (flat control) is the whole specificity story: Ouro's depth→time drift is GLOBAL, not language-specific. Faint vlines = recurrent-pass boundaries (s0|s1|s2|s3). Raw-argmax read-out, matched to the GPT-2 combined plot. Regenerated from best_run via analysis/fig_ouro_combined_scatter.py.
Gold standard — papermatch (clean regression-line Fig 3, Ouro)
Ouro in the same gold-standard papermatch format as the GPT-2 panel: clean single-row scatter, bold dashed regression line, big bold r per panel. IFG +0.66 · aSTG +0.60 (sig) · TP +0.25 (ns, n=6 underpowered) · mSTG-proxy +0.84. x-axis = effective recurrent depth (0→191); faint vertical lines = recurrent-pass boundaries (s0|s1|s2|s3). NOTE: the real EAC auditory control shows this is not language-specific — see poster v2. Regenerated from best_run via analysis/fig_papermatch_ouro.py.
Top row — per-representation curves + peak-lag dots (color = effective recurrent depth 0→191)
48 sampled representations across the 4 unrolled passes, same electrodes/estimator as the GPT-2 replication.
Scatter — peak lag vs effective recurrent depth (vertical lines = pass boundaries s0|s1|s2|s3)
IFG +0.66 · aSTG +0.60 (significant) · TP +0.25 (ns, n=6, saturates) · mSTG control +0.84 (significant — NOT flat). Iteration tracks brain time in language ROIs, but the control is not yet flat → specificity is the open question.

4 · Best validated run (recreated + verified)

Both encodings re-run from scratch and verified bit-exact vs the frozen oracle (max abs scores diff = 0.0, fingerprints match ORACLE.md). All 8 ROI numbers reproduce within ±0.01. These panels regenerate from best_run/ (RLM_DATA_ROOT=…/best_run). See best_run/VALIDATION.md.

GPT-2 XL — Goldstein Fig 3 (papermatch), verified
TP +0.89 · IFG +0.83 · aSTG +0.81 · mSTG −0.26 (ns) — every ROI exactly on oracle (Δ=0.00). Scores bit-identical to frozen (max abs diff 0.0); fingerprint 310abf2f… matches ORACLE.md (441 files). Regenerates from best_run/ via analysis/fig_papermatch.py + best_run/fig3_gpt2_best.py.
Ouro-2.6B — Goldstein Fig 3 (predshift), verified
TP +0.25 (ns) · IFG +0.66 · aSTG +0.60 · mSTG +0.84 — every ROI exactly on oracle (Δ=0.00). Scores bit-identical to frozen (max abs diff 0.0); fingerprint 3223fce4… matches ORACLE.md (432 files). Regenerates from best_run/ via analysis/fig_ouro_fig3.py + best_run/assemble_fig3_ouro.py.

5 · New poster (data-chosen, validated)

The data-chosen poster (SPEC §7): headline = whichever surviving claim cleared the full rigor bar + an adversarial skeptic. CPU-only re-analysis on frozen, cached encodings — no GPU run; the true EAC/Heschl auditory control (results_ouro_eac, 9 subj, 105 electrodes) replaces the old posterior-STG proxy; peak read-out unified to sub-bin parabolic; TP (n=6) labeled underpowered; the step-vs-layer panel is exploratory/within-model. A dense 192-rep re-confirmation is pending (deferred GPU job). Poster: https://uriubuntuserver.tail118306.ts.net/files/poster_v2.html

HEADLINE (survived rigor + skeptic) — per-ROI Ouro−GPT-2 peak-encoding gain, language-specific
On bit-identical electrodes (315 language, 105 EAC), Ouro-2.6B's recurrence beats GPT-2 XL's depth in peak encoding (max over reps × lags, 0–1000 ms) in higher-order language cortex — IFG +0.031 r (Ouro wins 85% of electrodes, 5/5 subjects, FDR q=1.5e−7), aSTG +0.034 r (86%, 7/7 subjects, q=2.8e−7) — and is ABSENT in the true-auditory EAC control (−0.013 r, Ouro wins only 38%). TP (n=6) is degenerate/underpowered (ns), never headlined. Survives bootstrap CIs over electrodes AND subjects, a label-swap permutation null, a window×read-out×pool-size sweep (0 sign-flips), and a mean-of-pool skeptic with NO best-of-pool selection (IFG +0.030 p=2.4e−7, aSTG +0.027 p=2.3e−6). Source: poster_v2/results/h3_per_roi.csv + figure_spec.
FOUNDATION — fair-pool global win + dimensionality controls (survived)
Pooled over 114 GPT-2-significant language electrodes (8 subj), Ouro (48 states) beats GPT-2 XL (49 layers) in peak encoding r: +0.030 mean gain, 85.1% of electrodes above y=x, Wilcoxon p=3.0e−13 (q_FDR=1.5e−12), Cohen's dz=+0.85, 8/8 subjects positive. Not a pool-size artifact (49→48 drop-layer-0 + ×200 subsample unchanged), not a feature-dim artifact (dim-insensitive median-rep still +0.022, p=1.4e−10; the EAC control reverses to GPT-2's favour −0.013). 7/7 validations pass, 0 sign-flips over 21 robustness cells; two independent code paths agree bitwise (r=1.0). Source: poster_v2/results/h1_full.json + figure_spec.
EXPLORATORY (within-model) — recursion step vs within-pass layer on brain latency
Decomposing Ouro's effective depth (eff = step×48 + layer) into its two axes: the across-pass RECURSION STEP (S 0–3) drives brain peak-latency in language cortex (aSTG +28 ms/pass r=0.97, IFG +10 ms/pass r=1.00; perm-FDR q<0.001) while the within-pass LAYER is flat. EXPLORATORY / within-model: an adversarial skeptic shows the EAC auditory control ALSO drifts with step (slope exceeds language ROIs), so this is presented as a within-model mechanism, NOT a language-specific claim — that role is carried by the headline figure. TP (n=6) hatched/degenerate; EAC step bar flagged read-out noise. Source: poster_v2/results/centerpiece_summary.csv + figure_spec.
Assembled poster (self-contained HTML → PNG/PDF)
Full poster: data-chosen headline (looping out-encodes depth, language-specifically) + the recursion-vs-depth centerpiece (exploratory) + the GPT-2 foundation + the real EAC specificity control. CIs + n on every panel, colorblind-safe (Okabe-Ito), self-contained captions, exploratory/underpowered cells labeled. Live: https://uriubuntuserver.tail118306.ts.net/files/poster_v2.html (PDF: /files/pv2_poster.pdf).

6 · Null-model (random-weight Ouro)

Is the depth→time drift INTRINSIC to recurrence dynamics, or a property of the TRAINED model? We re-encoded a random-weight (untrained) looped Ouro on bit-identical electrodes/ROIs/estimator and compared its drift to the trained model's. If the drift were architectural, random would drift too — including the GLOBAL EAC auditory control. It does not.

TRAINED vs RANDOM — peak lag vs effective recurrent depth (5 ROIs × 2 models)
VERDICT: TRAINING-DEPENDENT — the hypothesis FAILS. The untrained (random-weight) Ouro does NOT reproduce the depth→time drift. In the trained model the drift is robust and GLOBAL (IFG r +0.66 slope +0.18 ms/rep, aSTG +0.60/+0.55, mSTG-proxy +0.84/+0.79, and the true-auditory EAC control rises strongly +0.61/+3.35 ms/rep — slope-CIs exclude 0 over electrodes AND subjects). In the random model EVERY ROI's slope-CI straddles 0 (all ns: IFG r +0.34, aSTG +0.55, TP +0.15, mSTG +0.09) and the EAC control goes FLAT/NEGATIVE (r −0.23, slope −1.18 ms/rep) — opposite sign to trained. Random absolute encoding r collapses to ~0.10 (near-zero, expected for random weights) yet the peak-lags are non-degenerate (9–11 distinct lags, low modal fraction), so the random flatness is genuine noise, NOT a pinned-read-out artifact. The depth→time drift is a property of the TRAINED recurrence dynamics, not intrinsic to the looped architecture. Source: poster_v2/results/nullmodel_results.json + nullmodel_table.csv; regen via poster_v2/analysis/nullmodel_analysis.py + nullmodel_figure.py.

7 · Fig 3 — raw (un-normalized) encoding curves

The SAME Fig 3 analysis, but the TOP-row curves are NOT peak-normalized — each curve plots its RAW mean encoding r, so amplitude differences between layers/representations are visible (which layers/passes encode the brain strongest, not just where each peaks). Peak dots sit ON each curve at its actual (peak_lag, peak_value). Bottom-row peak-lag-vs-depth scatter + r is UNCHANGED (peak-lag is independent of curve normalization). Regenerate from best_run via analysis/fig3_unnormalized_gpt2.py and analysis/fig3_unnormalized_ouro.py.

GPT-2 XL — raw (un-normalized) encoding curves + peak-lag scatter
Top row is now TRUE encoding amplitude (y = encoding r, ~0.03→0.22), not peak-normalized to 1: early layers (blue) encode weakest, mid-to-late layers (green→red) strongest, with peak dots sitting on each raw curve. Bottom-row regression r unchanged — TP +0.89 · IFG +0.83 · aSTG +0.81 · mSTG −0.26 (ns, posterior-STG control). predshift, window 0–1000 ms. Regenerated from best_run via analysis/fig3_unnormalized_gpt2.py.
Ouro-2.6B — raw (un-normalized) encoding curves + peak-lag scatter
Top row is TRUE encoding amplitude (y = encoding r, ~0.02→0.27) over the 48 sampled representations, colored by effective recurrent depth 0→191; the very first representation (depth 0, dark blue) encodes markedly weakest, later passes encode strongest — invisible when every curve is normalized to 1. Peak dots sit on each raw curve. Bottom-row regression r unchanged — IFG +0.66 · aSTG +0.60 (sig) · TP +0.25 (ns) · mSTG-proxy +0.84; vlines = recurrent-pass boundaries s0|s1|s2|s3. Regenerated from best_run via analysis/fig3_unnormalized_ouro.py.