Failed Is Not a Terminal Status: A Contract for Agents

October 6, 2026AI & Automation9 min read
Failed Is Not a Terminal Status: A Contract for Agents
TL;DR: An agent that polls a slow, paid API until it reads failed will stop at the first one. In our API a run can fail because we stopped waiting, while the paid work behind it can still finish, so for that one case we made failed reopenable. The general rule: a status is a contract, the contract lives in the tool description, and failed needs a reason before an agent may act on it as final.

You wrote the loop the docs describe. Poll the job until it reaches a terminal status, then act on it: succeeded, fetch the file; failed, stop and report. Replicate's docs(opens in new tab) put it plainly: poll "until the prediction is in a terminal state (succeeded or failed)." Its lifecycle page adds canceled and aborted, and we come back to why.

For most jobs that loop is right. For a slow, paid job that someone else runs, it has one hole, and an agent walks into it faster than you would. This post will not tell you how long to wait, because the contract it describes does not say.

The same gap, in three other places

Three public reports show a consumer holding a status that does not match what happened to the job, or a status with no cause attached.

On 2026-03-01 a collaborator on the 3DStreet(opens in new tab) project opened an issue about a Pro user who generated two ten-second AI videos. Both generations succeeded on the hosting platform. The user saw "processing never ended", the videos were never saved, and 40 tokens (20 each) were deducted. The issue's author traces it to a server call that blocked for minutes and a client connection that dropped before the response arrived. Their stopgap, in their words: "Temporary solution for now -- make faster models the default."

GitHub's runner is a second case. An issue filed in March 2026(opens in new tab) describes ephemeral runners that finished a job with result Succeeded and were then reported as "lost communication with the server". The reporter says that in some cases the workflow_job.completed webhook never arrives, and puts the rate at 5 to 10 percent of their runs.

The third is a status object with no reason in it. In a Sora 2 API thread(opens in new tab), users posted jobs that read "status": "failed", "progress": 100, "error": null. An agent that reads that object has a word and a number that disagree, and nothing to branch on.

A human who sees any of these goes and looks. An agent does not, unless the contract tells it to.

How a run fails while the work can still finish

Clipwright(opens in new tab) is our video API for agents. A make_ugc run hands the slow part to a third-party vendor, and we wait for the result. Then we stop waiting. Our own timeout ends and the run is marked failed. The vendor job does not know about our timeout and may carry on. If it delivers, there is a paid video and a run that says failed.

In release 0.19.0 (2026-09-22) we changed the contract for that one case: a make_ugc run that failed because we stopped waiting for the vendor. If its paid vendor job is still held, the run can be reopened. It goes from failed back to queued and may reach succeeded. One run is reopened at most three times, and every reopening is named in warnings[]. Every other cause of failed is final.

The run contract: failed can lead back to queued queued in progress succeeded failed delivered we stop waiting reopened, at most 3 times, each named in warnings[] the contract, not measured data

The contract for a make_ugc run that failed because we stopped waiting. The arrow back to queued is the change.

The changelog states the consequence in one sentence: an agent that stops polling the moment it reads failed will miss a video that was paid for and delivered. It marks the change Breaking.

For a sense of what a run delivers, here is one finished Clipwright video: 28 seconds, made on production from a 70-word script. It is a faceless video, a different kind of run from the make_ugc one above.

A 28-second faceless video made by Clipwright on production

A status is a promise, and the tool description is where the agent reads it

The server side of this was live from the day it was deployed. Release 0.19.0 of the MCP server carried the contract change into the tool description.

That order is the part worth taking. An agent does not read your server code, and it does not read your changelog. It reads the sentence in the tool description and believes it. A behaviour that exists on the server but not in that sentence is, for the agent, not part of the contract. If the sentence still says a status is final, the agent treats it as final, whatever the server later does.

Which failed is this one?

Today the reason a failed run may reopen arrives as prose in warnings[]. That is the same array described in warnings are for humans, refusals are for agents: strings written for a person, with no code, that change when behaviour changes. So read what follows as what a run says today, not as values to match.

A run that failed because we stopped waiting says, in these words today, that "we stopped waiting while the vendor was still processing the video" and that "its paid job reference was retained for recovery". A run whose terms do not allow late recovery says the reference "is retained, but this run's terms do not support late recovery". A reopened run says "this run was reopened to collect a paid vendor job that landed late".

Table: Situation, What it promises, What the agent should do
SituationWhat it promisesWhat the agent should do
Failed because we stopped waiting, paid job reference retained for recovery
The run can be reopenedDo not treat it as final. Check the run again later
Failed, reference retained, but this run's terms do not support late recovery
No reopening under this run's termsTreat it as final
Failed for any other cause
Nothing followsTreat it as final

When a run's terms could not be read, it promises nothing either way, so plan around neither outcome.

There is no time limit in this contract, and we promise none. So the advice is not "wait". It is: when a failed run says its paid job was kept for recovery, do not treat it as final, and check the run again later instead of holding a polling loop open.

An agent that reads the sentences can act on their meaning. A script that matches the exact strings will break the day we reword one. That is the price of the decision the earlier post describes: we kept warnings as prose and refused to give them codes we could not keep, and the cost lands here, where a reason that decides whether an agent keeps waiting lives in a sentence. In your own API, put the difference between recoverable and final in a machine-readable field, and write down what each value allows next.

What it costs, and what it does not

On the bill, our rule from the bill is authorised by delivery is that a run failing after the vendor took the job is charged, unless the loss is ours. Our own wait running out counts as ours, so the failed run is charged nothing. If the run is reopened and the video arrives, it is charged once, by its measured length, like any other delivery. There is never a second bill. Four outcomes, not two has the longer argument.

The honest cost is on the agent's side. A status that can reopen is a surprise state transition, and not every system should allow one. A Cylc developer wrote in a 2021 issue(opens in new tab) about polling that returned a failed task to running: "It can never be safe given that the task has already entered the failed state and potentially triggered other tasks as a result."

They are right about unscoped resurrection. Anything downstream of a failed may already have acted on it. Our answer is to scope it until the surprise is predictable: one cause, three reopenings at most, every reopening named. We do not claim that makes it free. An agent that built a fallback on the first failed still has to decide what to do when the original arrives.

Other APIs reach for a different tool: more statuses. Replicate separates canceled from aborted(opens in new tab) by whether a prediction had started, and bills them differently. OpenAI's Batch API has an expired status(opens in new tab) that returns completed work and charges for it. Both put the cause in the status word. That works when the causes are few. When they are many, a reason beside the status scales better than a new status for each cause.

Check your own

Open the tool description for the slowest paid call your agent makes and find the sentence about failed. If it does not say what can follow it, add that sentence before you change anything else.

FAQ

Can a failed run really become succeeded?

In Clipwright, in one case. A make_ugc run that failed because we stopped waiting for the vendor, while the paid vendor job is still held, can go back to queued and reach succeeded. Every other cause of failed is final.

How many times can a run be reopened?

At most three. Each reopening is named in the run's warnings[].

Is there a time limit for a reopened run?

The contract has none, and we promise none. Check the run again later rather than holding a polling loop open.

What does a recovered delivery cost?

The failed run is charged nothing. A recovered delivery is charged once, as an ordinary one, by measured length. There is never a second bill.

Sources

Your agent stops at the first "failed". Is that word final?

A loop that stops at the first "failed" can miss a video that was paid for and delivered. In Clipwright a make_ugc run that failed because we stopped waiting can be reopened, and the tool says so.

About the Author

Dzmitry Vladyka
Dzmitry Vladyka

Dimantika

Co-founder of Dimantika. Builds Clipwright and ViralFaceless with coding agents. Previously ran GlockSoft with a partner for about 15 years. Writes about products and finding customers.

View all posts