17 August 2026

Evaluation tools shift focus from single models to full systems

  • New tools like eval-skills and Agent Arena measure how AI systems actually perform in real workflows, not just how well individual models score on tests.
  • These tools track practical concerns: whether systems route questions correctly, break problems into steps, remember context, and verify their own answers.
  • The shift matters because a great model inside a poorly designed system produces worse results than a mediocre model in a well-built one.

How it was covered

Latent Spaceswyx & Alessio

Tools like Hamel Husain's eval-skills plugin and Agent Arena's filters are shifting evaluation from model-level evals to harness-level measurement covering routing, decomposition, memory, verifier loops, and total completion cost. The newsletter notes this represents movement toward measuring full system behavior rather than isolated model performance.