🧪 Introducing Evals Measure Hex Agent performance and test context changes safely before t
🧪 Introducing Evals Measure Hex Agent performance and test context changes safely before they ship, all from the CLI. If you've added a guide or tweaked a semantic model, Evals helps you measure whether it actually helped — before you publish. To do that, define a suite of test questions with expect