Source record · content verified
AgentBench: Evaluating LLMs as Agents
benchmark documentation2023
Why this source is here
Multi-environment agent evaluation, cited for evaluation design rather than rankings.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Liu et al.
- Identifier
- arXiv:2308.03688