Failed Is Not a Terminal Status: A Contract for Agents

TL;DR: An agent that polls a slow, paid API until it readsfailedwill stop at the first one. In our API a run can fail because we stopped waiting, while the paid work behind it can still finish, so for that one case we madefailedreopenable. The general rule: a status is a contract, the contract lives in the tool description, andfailedneeds a reason before an agent may act on it as final.
You wrote the loop the docs describe. Poll the job until it reaches a terminal status, then act on it: succeeded, fetch the file; failed, stop and report. Replicate's docs(opens in new tab) put it plainly: poll "until the prediction is in a terminal state (succeeded or failed)." Its lifecycle page adds canceled and aborted, and we come back to why.
For most jobs that loop is right. For a slow, paid job that someone else runs, it has one hole, and an agent walks into it faster than you would. This post will not tell you how long to wait, because the contract it describes does not say.
The same gap, in three other places
Three public reports show a consumer holding a status that does not match what happened to the job, or a status with no cause attached.
On 2026-03-01 a collaborator on the 3DStreet(opens in new tab) project opened an issue about a Pro user who generated two ten-second AI videos. Both generations succeeded on the hosting platform. The user saw "processing never ended", the videos were never saved, and 40 tokens (20 each) were deducted. The issue's author traces it to a server call that blocked for minutes and a client connection that dropped before the response arrived. Their stopgap, in their words: "Temporary solution for now -- make faster models the default."
GitHub's runner is a second case. An issue filed in March 2026(opens in new tab) describes ephemeral runners that finished a job with result Succeeded and were then reported as "lost communication with the server". The reporter says that in some cases the workflow_job.completed webhook never arrives, and puts the rate at 5 to 10 percent of their runs.
The third is a status object with no reason in it. In a Sora 2 API thread(opens in new tab), users posted jobs that read "status": "failed", "progress": 100, "error": null. An agent that reads that object has a word and a number that disagree, and nothing to branch on.
A human who sees any of these goes and looks. An agent does not, unless the contract tells it to.
How a run fails while the work can still finish
Clipwright(opens in new tab) is our video API for agents. A make_ugc run hands the slow part to a third-party vendor, and we wait for the result. Then we stop waiting. Our own timeout ends and the run is marked failed. The vendor job does not know about our timeout and may carry on. If it delivers, there is a paid video and a run that says failed.
In release 0.19.0 (2026-09-22) we changed the contract for that one case: a make_ugc run that failed because we stopped waiting for the vendor. If its paid vendor job is still held, the run can be reopened. It goes from failed back to queued and may reach succeeded. One run is reopened at most three times, and every reopening is named in warnings[]. Every other cause of failed is final.
The contract for a make_ugc run that failed because we stopped waiting. The arrow back to queued is the change.
The changelog states the consequence in one sentence: an agent that stops polling the moment it reads failed will miss a video that was paid for and delivered. It marks the change Breaking.
For a sense of what a run delivers, here is one finished Clipwright video: 28 seconds, made on production from a 70-word script. It is a faceless video, a different kind of run from the make_ugc one above.
A status is a promise, and the tool description is where the agent reads it
The server side of this was live from the day it was deployed. Release 0.19.0 of the MCP server carried the contract change into the tool description.
That order is the part worth taking. An agent does not read your server code, and it does not read your changelog. It reads the sentence in the tool description and believes it. A behaviour that exists on the server but not in that sentence is, for the agent, not part of the contract. If the sentence still says a status is final, the agent treats it as final, whatever the server later does.
Which failed is this one?
Today the reason a failed run may reopen arrives as prose in warnings[]. That is the same array described in warnings are for humans, refusals are for agents: strings written for a person, with no code, that change when behaviour changes. So read what follows as what a run says today, not as values to match.
A run that failed because we stopped waiting says, in these words today, that "we stopped waiting while the vendor was still processing the video" and that "its paid job reference was retained for recovery". A run whose terms do not allow late recovery says the reference "is retained, but this run's terms do not support late recovery". A reopened run says "this run was reopened to collect a paid vendor job that landed late".
| Situation | What it promises | What the agent should do |
|---|---|---|
Failed because we stopped waiting, paid job reference retained for recovery | The run can be reopened | Do not treat it as final. Check the run again later |
Failed, reference retained, but this run's terms do not support late recovery | No reopening under this run's terms | Treat it as final |
Failed for any other cause | Nothing follows | Treat it as final |
When a run's terms could not be read, it promises nothing either way, so plan around neither outcome.
There is no time limit in this contract, and we promise none. So the advice is not "wait". It is: when a failed run says its paid job was kept for recovery, do not treat it as final, and check the run again later instead of holding a polling loop open.
An agent that reads the sentences can act on their meaning. A script that matches the exact strings will break the day we reword one. That is the price of the decision the earlier post describes: we kept warnings as prose and refused to give them codes we could not keep, and the cost lands here, where a reason that decides whether an agent keeps waiting lives in a sentence. In your own API, put the difference between recoverable and final in a machine-readable field, and write down what each value allows next.
What it costs, and what it does not
On the bill, our rule from the bill is authorised by delivery is that a run failing after the vendor took the job is charged, unless the loss is ours. Our own wait running out counts as ours, so the failed run is charged nothing. If the run is reopened and the video arrives, it is charged once, by its measured length, like any other delivery. There is never a second bill. Four outcomes, not two has the longer argument.
The honest cost is on the agent's side. A status that can reopen is a surprise state transition, and not every system should allow one. A Cylc developer wrote in a 2021 issue(opens in new tab) about polling that returned a failed task to running: "It can never be safe given that the task has already entered the failed state and potentially triggered other tasks as a result."
They are right about unscoped resurrection. Anything downstream of a failed may already have acted on it. Our answer is to scope it until the surprise is predictable: one cause, three reopenings at most, every reopening named. We do not claim that makes it free. An agent that built a fallback on the first failed still has to decide what to do when the original arrives.
Other APIs reach for a different tool: more statuses. Replicate separates canceled from aborted(opens in new tab) by whether a prediction had started, and bills them differently. OpenAI's Batch API has an expired status(opens in new tab) that returns completed work and charges for it. Both put the cause in the status word. That works when the causes are few. When they are many, a reason beside the status scales better than a new status for each cause.
Check your own
Open the tool description for the slowest paid call your agent makes and find the sentence about failed. If it does not say what can follow it, add that sentence before you change anything else.
FAQ
Can a failed run really become succeeded?
In Clipwright, in one case. A make_ugc run that failed because we stopped waiting for the vendor, while the paid vendor job is still held, can go back to queued and reach succeeded. Every other cause of failed is final.
How many times can a run be reopened?
At most three. Each reopening is named in the run's warnings[].
Is there a time limit for a reopened run?
The contract has none, and we promise none. Check the run again later rather than holding a polling loop open.
What does a recovered delivery cost?
The failed run is charged nothing. A recovered delivery is charged once, as an ordinary one, by measured length. There is never a second bill.
Sources
- Clipwright(opens in new tab), release 0.19.0 (2026-09-22) of the MCP server, from its changelog
- Replicate: create a prediction(opens in new tab) and prediction lifecycle(opens in new tab)
- OpenAI Batch API FAQ(opens in new tab)
- 3DStreet issue 1469(opens in new tab), actions/runner issue 4309(opens in new tab), Cylc issue 4513(opens in new tab)
- Sora 2 API thread, OpenAI Developer Community(opens in new tab)
Your agent stops at the first "failed". Is that word final?
A loop that stops at the first "failed" can miss a video that was paid for and delivered. In Clipwright a make_ugc run that failed because we stopped waiting can be reopened, and the tool says so.
About the Author
Dimantika
Co-founder of Dimantika. Builds Clipwright and ViralFaceless with coding agents. Previously ran GlockSoft with a partner for about 15 years. Writes about products and finding customers.
View all postsRelated posts
More articles you might like.

A New Enum Value Breaks Clients That Trusted Your List
A new enum value passes every schema diff and still breaks clients that validate responses. Read as text, send from a closed list, and mark it Breaking.

I Used to Handle Marketing. Now I Manage Coding Agents Too.
Coding agents have changed my working day. I still have to decide who the products are for, and I want more than my own assumptions to work with.

What One Second of AI Video Costs Us, Layer by Layer
One clip passes through five paid layers. We measured one of them with a wallet. The rest are price lists and estimates, and that gap is the honest answer.