We hold our AI agents to a higher standard than we hold ourselves.
Last month it took me and Claude nine drafts to write a single Slack reply. It was an escalation. A team wanted something from us that wasn’t going to happen this quarter, and I needed to say no without sounding like a jerk. The first draft read like a corporate memo. The second offered to size work we weren’t doing, and the third offered to join a client call I had never agreed to. By the ninth, I’d spent more time on the reply than on the problem it was about.
What makes it funny in hindsight is that two hours earlier I’d been in a product review arguing about evals for our tax agents. We spent most of that meeting on what belongs in the golden set and what error rate we’re willing to accept before a K-1 goes out to 400 investors. When you build software for real estate investment managers, a wrong number ends up on somebody’s tax return, so we take that conversation very seriously.
If one of our agents had a 3% error rate, we’d open a war room. Me rewriting the same reply nine times got a shrug and another coffee. That gap bothered me enough that I started running evals on myself.
Turns out it’s everyone
When I went looking for data, I found three numbers that made me feel better and worse at the same time. Stanford and BetterUp found that most “workslop,” AI output that looks polished and falls apart on a second read, takes about two hours to clean up. Workday blames “one-size-fits-all” tools for the rework that eats the time AI saves. The one that hurt came from Microsoft and Salesforce researchers: once a model makes a wrong assumption early in a conversation, it tends to double down.
That last finding describes my nine-draft reply almost exactly. Every “no, try again” makes the next answer less reliable, and you pay tokens for the privilege, since each retry resends the whole conversation.
Generic in, generic out
The root problem is context. An AI that knows nothing about you will write the average version of whatever you ask for. Ask it for a PRD and you’ll get one that could belong to any company on earth, complete with a persona named “Sarah, 34, busy professional” and a success metric of “increase engagement.” Then you spend the next hour turning that average document into yours by explaining your customer, then your format, then the fact that you hate the word “seamless.” The next morning you open a new chat and do it all over again, because the chat has forgotten everything.
Treating my workflow like a product
The fix I landed on was to take the same playbook we use for our agents and point it at myself:
- Flag on the second redo. Once is noise. Twice is a bug.
- Log it in one line. What I asked for and what I got, plus the part that matters: what I actually wanted.
- Fix the source. Put the fix somewhere the AI reads every session, never just in the chat.
- Retest. Open a fresh session, replay the original request, and see if the failure comes back.
The flagging part runs itself now, because I added this rule to my instructions file. Feel free to steal it, since it’s probably the cheapest eval you’ll ever set up.
It fired while I was writing this post, after I rejected a headline for being clickbait:
Where the fix goes
Nearly every fix I’ve logged lands in one of three places. Choosing the wrong one is how you end up with a 4,000-word prompt that nobody can maintain, including you.
Here’s what each one looks like in practice.
Instructions. For a while, the AI kept inventing a brand-new structure for documents that were supposed to replace an existing one. I got three wrong versions in a single afternoon before I realized the fix was one line:
- If the source document says it replaces or supersedes a prior artifact, read that artifact before proposing any structure.
Skill. The nine-draft reply became a skill called cooldown. Before it writes anything, it asks me two questions:
- How hard is this no: not now with no date, or is there a real path?
- Is there anything you’re actually willing to offer, or is the answer just no?
It also has one hard rule, which is to never invent a concession I didn’t offer. That rule exists because drafts two, three and seven each offered something I never agreed to give.
Knowledge base. The AI also had a habit of stating IRS K-1 box locations from memory, confidently and sometimes wrongly. In tax, that kind of mistake ends up on an investor’s form, so the fix was a K-1 reference in my knowledge base plus a rule to go with it:
- Cite the knowledge base, or mark the field “verify against current form.” Never state recalled form structure as fact.
Looking back across my log, roughly half the fixes went into instructions, about a third became skills, and the rest went into the knowledge base.
It gets quieter
The obvious worry is that all of this turns into a second job, but my log says otherwise.
May was the first month I ran the loop for real, and everything surfaced at once. Logging 45 failures was mildly humbling. After that, new failures dropped about 75% and stayed there. September ticked up again when I moved into new kinds of work, which makes sense, because new work brings new failures. The old ones stayed dead.
What I didn’t expect was how much the log itself would change. What I’m logging today looks nothing like May. The basics are handled, so what I’m teaching the AI now is harder and much more specific to tax and real estate, and I’ve stopped re-explaining who I am every morning.
Start in fifteen minutes
You don’t have to wait a week to collect data. You already have it.
- Retro last week. Scroll back through last week’s chats and find every place you typed “no,” “not quite,” or “try again.” Paste each exchange into a note called “things I keep correcting.”
- Hand the note back to the AI with this: “Group these by the underlying mistake. For each group, tell me whether the fix belongs in my instructions, a skill, or my knowledge base, and draft it.”
- Fix the repeats today. Anything that showed up twice gets fixed now.
- Paste the flag rule into your instructions so it catches repeats from now on.
- Repeat on Friday. Run the same retro on this week. A month in, check whether the old lines stopped showing up.
What I learned
The biggest lesson for me is that the second time you correct the AI, you’ve found a bug, and it’s worth logging right then. The chat forgets everything, but your setup doesn’t, so the fix has to live somewhere the AI will read next time. How you work goes in instructions, repeated tasks become skills, and facts go in the knowledge base. If you try this, expect the first month to be loud. Mine was, and it got quiet fast.
We make our agents learn from their mistakes before they touch a customer, and it seems fair to ask the same of ourselves.
The second correction is free data. The tenth is a choice.





