
I Used to Handle Marketing. Now I Manage Coding Agents Too.
Coding agents have changed my working day. I still have to decide who the products are for, and I want more than my own assumptions to work with.
24 posts

Coding agents have changed my working day. I still have to decide who the products are for, and I want more than my own assumptions to work with.

Paying a vendor feels like it should authorise charging the client. It does not, and the wording that says it does once billed 900 credits for a video that did not exist.

Three products now let the user's own Claude plan pay for inference. Anthropic's terms permit one shape of that and ban another, and you cannot tell which.

No /openapi.json, no generator, no spec to drift. The machine-readable entry point is the MCP server's tools/list — and that turned out to be the honest choice, not the lazy one.

A 3% aspect mismatch is a rounding error. A 40% one is a different video. The threshold between them decides whether you refuse before the bill or apologise after.

An unmet parameter should come back named. We learned why the hard way: two warnings went missing for weeks and every test stayed green.

Six blockers across four consecutive stories, and the two reviews agreed on one of them. Here is what running both actually cost us, and the rule we ended up with.

Refusals in our API carry stable codes. Warnings carry none, on purpose — and the bug that proved why cost us two warnings nobody noticed were missing.

Our billing rule said: failed run after a paid vendor call means we charge. It sounded fair. On a run that only paid for speech, it was false, and it cost a customer.

Teams blame MCP when agents fumble their tools. We shipped an MCP server for a video API and found the real defect: offering a field you answer with a 400.

TL;DR: A 95%-accurate agent step sounds safe, but ten steps land you near 60% and twenty near 36%. Multi-step chains multiply their error. Cut the chain, verify between steps, gate the risky actions.

TL;DR: A third or more of AI agent failures aren't crashes. They're agents reporting success for work that didn't happen. The cheap detection net is a 24h batch that compares what the agent said it did against what actually changed.