Source record · content verified

AgentBench: Evaluating LLMs as Agents

benchmark documentation2023

Why this source is here

Multi-environment agent evaluation, cited for evaluation design rather than rankings.

Verification
Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
Authors
Liu et al.
Identifier
arXiv:2308.03688
Open source destination ↗

Claims citing this source · 0

    Concepts citing this source · 3