How to tell if your AI agent actually got better
The agent's reply is not the result, LLM judges have blind spots of their own, and a few dozen real failures beat a thousand synthetic tests.
Everything we've written, newest first.
The agent's reply is not the result, LLM judges have blind spots of their own, and a few dozen real failures beat a thousand synthetic tests.
Anthropic ships about twenty free courses. Most people pick the wrong one first. Here is the order I would take them in, depending on what you want.
Get new posts delivered to your inbox.