datacurve-ai/deep-swe
🟢Worth installingNew discovery — enrichment pending(New discovery — enrichment pending)
Measuring frontier coding agents on original, long-horizon engineering tasks
Why this repo matters
- •Evaluates frontier coding agents on real-world, long-horizon software engineering tasks.
- •Provides a standardized benchmark with 113 tasks across multiple programming languages.
- •Uses a sandboxed environment with program-based verifiers for objective evaluation.
Key Metrics
- Stars: 1.6k
- Forks: 106
- Open Issues: 72
- Stars (7d delta): 0
- Stars (30d delta): 0
Scores & Metadata
- Trend Velocity Score: 0.00
- Opportunity Score: 35.00
- Confidence Score: 20.00
- Language: Python
- First Seen: 1 month ago