# Synthetic Intelligence Atlas A source-bounded map of large language models, agents, evaluation, and forecasting limits. Each public claim carries sources, a status, and limitations. Version: 1.0.0 Evidence cutoff: 2026-06-14 Last reviewed: 2026-07-27 Canonical URL: https://research.mahastrategies.com/atlas/synthetic-intelligence Claims JSON: https://research.mahastrategies.com/atlas/synthetic-intelligence/claims.json Sources JSON: https://research.mahastrategies.com/atlas/synthetic-intelligence/sources.json ## Editorial boundary This first public edition includes only established and active-research claims supported by content-verified public sources. Forecast scenarios, URL-only sources, local audit artifacts, benchmark leaderboards, specific pricing trajectories, and quantum computing are excluded. ## Public claims - si-001 [established]: The transformer architecture replaces recurrence and convolution with attention as the primary sequence-modelling mechanism. - si-002 [established]: Sufficiently large autoregressive language models perform tasks from instructions and examples supplied in context, without gradient updates. - si-003 [established]: Language-model loss follows empirical power-law relationships with training compute, parameter count, and dataset size over the studied ranges. - si-004 [active]: The compute-optimal allocation between parameters and training tokens has been revised in the literature, and deployed practice diverges from compute-optimal training. - si-005 [established]: Fine-tuning with human feedback substantially changes how closely a model follows instructions relative to its pre-trained base. - si-006 [established]: Prompting a model to produce intermediate reasoning steps changes its measured performance on multi-step problems. - si-007 [active]: Allocating additional computation at inference time can, in studied settings, improve results more than spending the equivalent compute on additional parameters. - si-008 [established]: Retrieval-augmented generation conditions outputs on documents retrieved at inference time rather than on parametric memory alone. - si-009 [established]: Language models can be trained or prompted to invoke external tools and incorporate returned results. - si-010 [established]: The design of the agent-computer interface materially changes measured coding-agent performance holding the underlying model fixed. - si-011 [established]: SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue. - si-012 [established]: MMLU and GPQA are multiple-choice benchmarks over defined subject distributions, with human baselines reported by their authors. - si-013 [established]: Benchmarks are revised by their owners in response to saturation and robustness problems, producing successor versions that are not directly comparable to the originals. - si-014 [active]: Contamination of evaluation data into training corpora is a measured problem that must be assessed per benchmark rather than assumed absent. - si-015 [active]: The use of general-purpose benchmarks as measures of general capability is contested in the literature on construct-validity grounds. - si-016 [established]: Multi-metric evaluation across scenarios has been proposed and implemented as an alternative to single headline scores. - si-017 [active]: The behaviour of a served model accessed through an API can change over time on fixed prompts. - si-018 [established]: METR defines a 50%-task-completion time horizon: the time humans typically take on tasks a model completes with 50% success. Its published measurement reports a horizon of around 50 minutes for one frontier model of that period, and a doubling of the horizon roughly every seven months since 2019, with a possible acceleration in 2024. - si-019 [active]: Increases in measured long-horizon performance are attributed in part to improved reliability and error recovery rather than to reasoning improvements alone. - si-020 [established]: Provider technical reports and model cards are self-reports by the organization that built the system, and may withhold architecture, data, and training details. - si-021 [active]: Aggregated expert forecasts of AI capability timelines have shifted substantially between survey rounds. - si-023 [established]: A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability. ## Sources - Attention Is All You Need (2017) — https://arxiv.org/abs/1706.03762 - Language Models are Few-Shot Learners (2020) — https://arxiv.org/abs/2005.14165 - Scaling Laws for Neural Language Models (2020) — https://arxiv.org/abs/2001.08361 - Training Compute-Optimal Large Language Models (2022) — https://arxiv.org/abs/2203.15556 - PaLM: Scaling Language Modeling with Pathways (2022) — https://arxiv.org/abs/2204.02311 - Training language models to follow instructions with human feedback (2022) — https://arxiv.org/abs/2203.02155 - Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) — https://arxiv.org/abs/2201.11903 - Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (2024) — https://arxiv.org/abs/2408.03314 - Training Verifiers to Solve Math Word Problems (2021) — https://arxiv.org/abs/2110.14168 - Solving Quantitative Reasoning Problems with Language Models (2022) — https://arxiv.org/abs/2206.14858 - Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — https://arxiv.org/abs/2005.11401 - Toolformer: Language Models Can Teach Themselves to Use Tools (2023) — https://arxiv.org/abs/2302.04761 - ReAct: Synergizing Reasoning and Acting in Language Models (2022) — https://arxiv.org/abs/2210.03629 - SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023) — https://arxiv.org/abs/2310.06770 - SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (2024) — https://arxiv.org/abs/2405.15793 - AgentBench: Evaluating LLMs as Agents (2023) — https://arxiv.org/abs/2308.03688 - Measuring AI Ability to Complete Long Software Tasks (2025) — https://arxiv.org/abs/2503.14499 - Measuring Massive Multitask Language Understanding (2020) — https://arxiv.org/abs/2009.03300 - MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (2024) — https://arxiv.org/abs/2406.01574 - GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023) — https://arxiv.org/abs/2311.12022 - On the Measure of Intelligence (2019) — https://arxiv.org/abs/1911.01547 - Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models (2022) — https://arxiv.org/abs/2206.04615 - Holistic Evaluation of Language Models (2022) — https://arxiv.org/abs/2211.09110 - AI and the Everything in the Whole Wide World Benchmark (2021) — https://arxiv.org/abs/2111.15366 - What Will it Take to Fix Benchmarking in Natural Language Understanding? (2021) — https://arxiv.org/abs/2104.02145 - NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark (2023) — https://arxiv.org/abs/2310.18018 - Investigating Data Contamination in Modern Benchmarks for Large Language Models (2023) — https://arxiv.org/abs/2311.09783 - How is ChatGPT's behavior changing over time? (2023) — https://arxiv.org/abs/2307.09009 - On the Opportunities and Risks of Foundation Models (2021) — https://arxiv.org/abs/2108.07258 - Model Cards for Model Reporting (2019) — https://arxiv.org/abs/1810.03993 - GPT-4 Technical Report (2023) — https://arxiv.org/abs/2303.08774 - The Llama 3 Herd of Models (2024) — https://arxiv.org/abs/2407.21783 - Thousands of AI Authors on the Future of AI (2024) — https://arxiv.org/abs/2401.02843