
We Shipped an Agent-Native API Without OpenAPI
No /openapi.json, no generator, no spec to drift. The machine-readable entry point is the MCP server's tools/list — and that turned out to be the honest choice, not the lazy one.
13 posts

No /openapi.json, no generator, no spec to drift. The machine-readable entry point is the MCP server's tools/list — and that turned out to be the honest choice, not the lazy one.

An unmet parameter should come back named. We learned why the hard way: two warnings went missing for weeks and every test stayed green.

Six blockers across four consecutive stories, and the two reviews agreed on one of them. Here is what running both actually cost us, and the rule we ended up with.

Refusals in our API carry stable codes. Warnings carry none, on purpose — and the bug that proved why cost us two warnings nobody noticed were missing.

Our render container reported 32 cores when it had 4. Remotion sized its defaults from that and OOMed. Five confident diagnoses died before the answer turned up in a version field nobody had thought to print.

Our billing rule said: failed run after a paid vendor call means we charge. It sounded fair. On a run that only paid for speech, it was false, and it cost a customer.

Teams blame MCP when agents fumble their tools. We shipped an MCP server for a video API and found the real defect: offering a field you answer with a 400.

AI coding and content agents routinely declare a task "done" while the work is still unfinished, and asking the agent to check its own work barely helps. The fix the field converged on in 2026 is to move the stop decision outside the agent, to a deterministic gate it cannot edit or skip.

A viral thread promises Stanford's STORM research method in four prompts with no setup. It quietly deletes the one thing that makes STORM research instead of confident guessing — retrieval. Here's why grounding is the whole point, and how we keep it.

Elvis Sun's loss-function development reframes long agent loops: optimize toward a target, not a spec. We mapped it onto our content pipeline. Here's what we found.

Project Glasswing found 10,000+ severe bugs with AI. Small SaaS teams need shorter patch loops, cleaner dependency inventory, and agent audit trails.

Anyone can vibe-code your app in an afternoon. Three moats still hold: distribution, network effects, and data partnerships. Here's which ones we're betting on.