
Engineering / 14 min read
Benchmarking the Agents That Write Our Code
See how we built a benchmark from our own git history to compare AI coding models, skill packs, and agent harnesses on our real codebase.
- System
- Fullscript platform
- Stack
- Production systems
- Scale
- 100K+ practitioners
- Status
- Production
- Focus
- Engineering



