AgentBench
by THUDM
Pythonpushed almost 2 years ago
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
AI summary
Agent evaluation platform
A benchmark suite for evaluating the ability of large language models to operate as autonomous agents in various environments
- stars
- 2.3K
- forks
- 167
- watching
- 28
- awesome list
- 1
Featured in 1 awesome list
Each link jumps to the spot where the list mentions AgentBench.