17 August 2026
AI evaluation tools shift focus from models to workflows
- Developers are building tools like eval-skills plugins and Agent Arena that measure how AI systems perform in real workflows, not just raw model capability.
- These tools track practical outcomes: whether the system routes requests correctly, breaks problems into steps, remembers context, and stays within budget, not just accuracy scores.
- This shift reflects the field recognizing that a powerful model alone does not guarantee good results in production systems.
How it was covered
Latent Spaceswyx & Alessio
Tools like Hamel Husain's eval-skills plugin and Agent Arena's filters show the field moving from model-level evals to harness-level measurement including routing, decomposition, memory, and total completion cost. The newsletter emphasises this represents a maturation toward practical measurement.