Attention Is All You Need
Vaswani et al. · 2017
The architecture paper underlying the transformer models this atlas is about.
Source record →Atlas library · 33 content-verified sources
Every public source was content-verified during this review. A source record explains its role; it does not claim that the Atlas endorses the paper’s argument or that the field agrees with it.
Vaswani et al. · 2017
The architecture paper underlying the transformer models this atlas is about.
Source record →Brown et al. · 2020
Introduced in-context few-shot learning as an emergent property of scale.
Source record →Kaplan et al. · 2020
The original power-law scaling result, and one side of the compute-allocation disagreement.
Source record →Hoffmann et al. · 2022
The Chinchilla revision of compute-optimal allocation, and the other side of that disagreement.
Source record →Chowdhery et al. · 2022
A large-scale training report cited for scaling practice, not for leaderboard position.
Source record →Ouyang et al. · 2022
The post-training method that separated instruction-following from raw pre-training.
Source record →Wei et al. · 2022
Established that intermediate reasoning steps at inference change measured performance.
Source record →Snell et al. · 2024
The inference-time compute trade-off against parameter scaling.
Source record →Cobbe et al. · 2021
Verifier-based selection over sampled solutions, an early inference-time compute method.
Source record →Lewkowycz et al. · 2022
Domain-targeted training on quantitative reasoning; cited for method, not for score comparisons.
Source record →Lewis et al. · 2020
The originating formulation of retrieval-augmented generation.
Source record →Schick et al. · 2023
Learned external tool invocation from the model side.
Source record →Yao et al. · 2022
The interleaved reason-and-act pattern underlying most agent loops.
Source record →Jimenez et al. · 2023
The benchmark definition for test-verified resolution of real repository issues.
Source record →Yang et al. · 2024
The load-bearing source for separating scaffolding capability from model capability.
Source record →Liu et al. · 2023
Multi-environment agent evaluation, cited for evaluation design rather than rankings.
Source record →Kwa et al. (METR) · 2025
The primary definition of the 50%-task-completion time horizon. Abstract read during this pass; it is the only verified source for any time-horizon figure in this map.
Source record →Hendrycks et al. · 2020
The MMLU benchmark definition.
Source record →Wang et al. · 2024
A benchmark revision motivated by saturation and robustness problems in the original.
Source record →Rein et al. · 2023
The GPQA definition and its stated human baselines.
Source record →Chollet · 2019
The skill-versus-generalization argument and the ARC formulation, central to construct validity here.
Source record →Srivastava et al. · 2022
A large collaborative benchmark explicitly concerned with extrapolating capability.
Source record →Liang et al. · 2022
The argument for multi-metric evaluation instead of a single headline number.
Source record →Raji et al. · 2021
The construct-validity critique of general-purpose benchmarks.
Source record →Bowman and Dahl · 2021
Names the structural failure modes of benchmark-driven progress measurement.
Source record →Sainz et al. · 2023
Contamination as a per-benchmark measurement problem rather than an anecdote.
Source record →Deng et al. · 2023
Empirical contamination investigation across widely used benchmarks.
Source record →Chen, Zaharia, Zou · 2023
Documents that a fixed prompt against a served model is not a stable measurement over time.
Source record →Bommasani et al. · 2021
Framing source for governance and accountability of general-purpose models.
Source record →Mitchell et al. · 2019
The reporting convention whose existence is why provider self-reports are structured the way they are.
Source record →OpenAI · 2023
Cited as an example of a provider technical report, including its own disclosure that architecture and training details are withheld. Not an independent measurement.
Source record →Grattafiori et al. (Meta) · 2024
A comparatively detailed provider training report; still a self-report.
Source record →Grace et al. · 2024
The large expert-survey instrument used for timeline forecasting, and evidence about forecast instability.
Source record →