THUDM/AgentBench
🟢Worth installingNew discovery — enrichment pending(New discovery — enrichment pending)
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Why this repo matters
- •Provides standardized evaluation framework for assessing LLM agent capabilities across diverse environments and tasks.
- •Accepted at ICLR 2024, establishing academic credibility for benchmarking autonomous language model agents.
- •Enables reproducible comparison of agent performance beyond simple chat completion metrics using Python tooling.
Key Metrics
- Stars: 3.7k
- Forks: 280
- Open Issues: 77
- Stars (7d delta): 0
- Stars (30d delta): 31
Scores & Metadata
- Trend Velocity Score: 0.00
- Opportunity Score: 25.00
- Confidence Score: 100.00
- Language: Python
- Topics:
- First Seen: 2 months ago
Star History
3.7k stars+89 over 10 snapshots8/8/2026 – 9/22/2026