Atlas library · 33 content-verified sources

A source trail you can inspect.

Every public source was content-verified during this review. A source record explains its role; it does not claim that the Atlas endorses the paper’s argument or that the field agrees with it.

primary paperContent verified

Attention Is All You Need

Vaswani et al. · 2017

The architecture paper underlying the transformer models this atlas is about.

Source record →
primary paperContent verified

Language Models are Few-Shot Learners

Brown et al. · 2020

Introduced in-context few-shot learning as an emergent property of scale.

Source record →
primary paperContent verified

Scaling Laws for Neural Language Models

Kaplan et al. · 2020

The original power-law scaling result, and one side of the compute-allocation disagreement.

Source record →
primary paperContent verified

Training Compute-Optimal Large Language Models

Hoffmann et al. · 2022

The Chinchilla revision of compute-optimal allocation, and the other side of that disagreement.

Source record →
primary paperContent verified

PaLM: Scaling Language Modeling with Pathways

Chowdhery et al. · 2022

A large-scale training report cited for scaling practice, not for leaderboard position.

Source record →
primary paperContent verified

Training language models to follow instructions with human feedback

Ouyang et al. · 2022

The post-training method that separated instruction-following from raw pre-training.

Source record →
primary paperContent verified

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Wei et al. · 2022

Established that intermediate reasoning steps at inference change measured performance.

Source record →
primary paperContent verified

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Snell et al. · 2024

The inference-time compute trade-off against parameter scaling.

Source record →
primary paperContent verified

Training Verifiers to Solve Math Word Problems

Cobbe et al. · 2021

Verifier-based selection over sampled solutions, an early inference-time compute method.

Source record →
primary paperContent verified

Solving Quantitative Reasoning Problems with Language Models

Lewkowycz et al. · 2022

Domain-targeted training on quantitative reasoning; cited for method, not for score comparisons.

Source record →
primary paperContent verified

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Lewis et al. · 2020

The originating formulation of retrieval-augmented generation.

Source record →
primary paperContent verified

Toolformer: Language Models Can Teach Themselves to Use Tools

Schick et al. · 2023

Learned external tool invocation from the model side.

Source record →
primary paperContent verified

ReAct: Synergizing Reasoning and Acting in Language Models

Yao et al. · 2022

The interleaved reason-and-act pattern underlying most agent loops.

Source record →
benchmark documentationContent verified

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Jimenez et al. · 2023

The benchmark definition for test-verified resolution of real repository issues.

Source record →
primary paperContent verified

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Yang et al. · 2024

The load-bearing source for separating scaffolding capability from model capability.

Source record →
benchmark documentationContent verified

AgentBench: Evaluating LLMs as Agents

Liu et al. · 2023

Multi-environment agent evaluation, cited for evaluation design rather than rankings.

Source record →
research organizationContent verified

Measuring AI Ability to Complete Long Software Tasks

Kwa et al. (METR) · 2025

The primary definition of the 50%-task-completion time horizon. Abstract read during this pass; it is the only verified source for any time-horizon figure in this map.

Source record →
benchmark documentationContent verified

Measuring Massive Multitask Language Understanding

Hendrycks et al. · 2020

The MMLU benchmark definition.

Source record →
benchmark documentationContent verified

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Wang et al. · 2024

A benchmark revision motivated by saturation and robustness problems in the original.

Source record →
benchmark documentationContent verified

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Rein et al. · 2023

The GPQA definition and its stated human baselines.

Source record →
primary paperContent verified

On the Measure of Intelligence

Chollet · 2019

The skill-versus-generalization argument and the ARC formulation, central to construct validity here.

Source record →
benchmark documentationContent verified

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Srivastava et al. · 2022

A large collaborative benchmark explicitly concerned with extrapolating capability.

Source record →
primary paperContent verified

Holistic Evaluation of Language Models

Liang et al. · 2022

The argument for multi-metric evaluation instead of a single headline number.

Source record →
primary paperContent verified

AI and the Everything in the Whole Wide World Benchmark

Raji et al. · 2021

The construct-validity critique of general-purpose benchmarks.

Source record →
primary paperContent verified

What Will it Take to Fix Benchmarking in Natural Language Understanding?

Bowman and Dahl · 2021

Names the structural failure modes of benchmark-driven progress measurement.

Source record →
primary paperContent verified

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Sainz et al. · 2023

Contamination as a per-benchmark measurement problem rather than an anecdote.

Source record →
primary paperContent verified

Investigating Data Contamination in Modern Benchmarks for Large Language Models

Deng et al. · 2023

Empirical contamination investigation across widely used benchmarks.

Source record →
primary paperContent verified

How is ChatGPT's behavior changing over time?

Chen, Zaharia, Zou · 2023

Documents that a fixed prompt against a served model is not a stable measurement over time.

Source record →
primary paperContent verified

On the Opportunities and Risks of Foundation Models

Bommasani et al. · 2021

Framing source for governance and accountability of general-purpose models.

Source record →
primary paperContent verified

Model Cards for Model Reporting

Mitchell et al. · 2019

The reporting convention whose existence is why provider self-reports are structured the way they are.

Source record →
provider self reportContent verified

GPT-4 Technical Report

OpenAI · 2023

Cited as an example of a provider technical report, including its own disclosure that architecture and training details are withheld. Not an independent measurement.

Source record →
provider self reportContent verified

The Llama 3 Herd of Models

Grattafiori et al. (Meta) · 2024

A comparatively detailed provider training report; still a self-report.

Source record →
primary paperContent verified

Thousands of AI Authors on the Future of AI

Grace et al. · 2024

The large expert-survey instrument used for timeline forecasting, and evidence about forecast instability.

Source record →