ninjahawk/livenerf
🟢Worth installingNew discovery — enrichment pending(New discovery — enrichment pending)
Benchmark for tracking model capability after release.
Why this repo matters
- •Tracks frontier model capability drift after release using a deterministic benchmark.
- •Employs a frozen prompt and pinned CLI approach to measure statistical drift.
- •Built on the UK AI Security Institute's Inspect eval framework for standardized testing.
Key Metrics
- Stars: 375
- Forks: 5
- Open Issues: 1
- Stars (7d delta): 0
- Stars (30d delta): 0
Scores & Metadata
- Trend Velocity Score: 0.00
- Opportunity Score: 50.00
- Confidence Score: 20.00
- Language: Python
- First Seen: 1 week ago