Skip to main content

Command Palette

Search for a command to run...

Series

Building an Agent Evaluation Platform

Documenting the build of an open-source evaluation framework for AI agents — starting with a multi-agent travel planner as the first system it grades. Real benchmarks, real failures, real fixes, across multiple models.