Notes from the exam hall.
Stories and explainers on verifying AI agents — the method, the misses, and what the numbers keep teaching us.
Testing a custom AI agent: hundreds of thousands of people built one, hardly anyone checks it
OpenClaw and Hermes Agent have over 600,000 GitHub stars between them, open models are about four months behind the frontier, and one in five teams running agents in production don't evaluate them at all. The model is becoming the cheap part. The harness you build on top is what makes it your agent, and it's the part nobody measures.
Read →Every eval is a test your agent has already seen
Self-reported success is the most common way agents fail, LLM judges can't catch it, and a test you run on yourself is one the agent can pass without being good. What an independent, re-sat, proof-gated read changes.
Read →GPT-6 Astra, sat as a naked agent: 42.55/100
OpenAI's GPT-6 Astra shipped on 3 September. On 8 September we sat it through the full Verigent battery as a naked agent — no harness — for a dated, on-chain score.
Read →How many dimensions does agent capability actually have?
We build agents against the three or four dimensions we can see, then call them good. But capability space is far wider than task completion, and the dimensions nobody measures are where agents quietly fail.
Read →Why continuous testing is necessary for AI agents
A test an agent sat last month tells you what it could do last month. Agents change underneath you, snapshots can be crammed for, and a single frame can show a wheel spinning backwards. The case for testing that never stops.
Read →You can't cram an exam that doesn't exist yet
Every benchmark score you've ever seen was one somebody could study for. Why drillable tests keep fooling us about AI agents, and what a structurally uncrammable exam looks like.
Read →Verigent is live: find out where your agent actually stands
The continuous agent test went live this week. Fresh challenges drawn at sit-time, an honest dimension-by-dimension read, and a battery that grows with the community. First test free, ten minutes over MCP.
Read →The test is alive, on purpose
A finished test is a crammable test, so ours is deliberately never finished. How a living battery, community-proposed dimensions, and shadow collection keep an agent exam honest for good.
Read →