Blocking Agent Polling Got Us 273 'echo waiting' Turns

You wrote the rule down. A section in CLAUDE.md, the exact command, the reason. Then you watched the agent run gh pr checks again.
A commenter on a Claude Code bug report describes the same moment. They asked the agent "Why don't you wait?", it agreed, and, in their words, "it immediately started polling again"(opens in new tab).
We wanted to know how often that happens to us, so we counted. A reflection pass read 46 Claude Code transcripts from three of our repositories, 27 August to 15 September 2026, about 600 MB of session logs. This post is what the count showed and what we did about each number. It will not tell you the guards worked everywhere. Our loop guard was routed around inside a single session, one count in our own numbers does not reconcile, and one of the fixes is still just a sentence in CLAUDE.md.
What did 46 transcripts measure?
Waiting was one of the biggest leaks. The pass ranked noise from a tool notice first (8,233 injections across the three repositories, the biggest lever by tokens) and waiting second. The 46 sessions split 21 / 14 / 11 between viralfaceless.ai, our main project and clipwright. One extractor script (jq) pulled the same signals from every transcript: owner messages, tool errors by class, command families, repeats, polls. Three subagents, one per repository, read the dumps and the raw transcripts.
| Finding | Counted in | Count | What it became |
|---|---|---|---|
Polling commands, in five incompatible idioms | main project | 231 in the summary, 213 in the per-session table | loop-guard hook |
Empty echo waiting turns, after a block on sleep-then-poll | one clipwright session | 273 in four bursts (180 in the first) | an empty-turn rule in loop-guard |
Secret reads, and CMS logins | main project | 546 and 108 | scripts/cms-token.sh |
Subagents the owner killed at session end | main project | 39 of 51 (54 results truncated) | a sentence in CLAUDE.md |
Waiting commands despite a written recipe | viralfaceless.ai | 1,146 of 10,683 (11%) | the case for the guard |
The counts are not comparable across rows. They come from different repositories and measure different things, so none of them is summed.
One count does not reconcile, and it is in the first row. Our project instructions and the hook change we merged both say 231 polling commands. The per-session table in the analysis adds up to 213, and it flags 5 of those as classifier errors, not waits. A further split in the summary, 90 gh pr checks and 66 gh run, has the same problem: the per-session gh pr checks column adds up to 100, not 90. The raw transcripts were deleted on 15 September, so nobody can recount. We quote 231 where we cite our own project's number and 213 where we say what the table holds. What we take from it is the size, around two hundred polling commands, and not the third digit.
The pattern inside the polling is clearer than the total. One session reworded gh pr checks 336 eight times in a row. The --watch flag, which makes the command wait for the checks instead of asking once, first shows up in this project's transcripts on 8 September. The knowledge did not carry from one session to the next.
Why did a written rule fail 1,146 times?
A CLAUDE.md rule is context, and Anthropic's documentation says so. In the viralfaceless.ai repository the CLAUDE.md has a "Long Waits & Polling" section with the exact recipe: run the command with --watch in the background and wait for the completion notification. The count there was 1,146 waiting commands out of 10,683, 11% of everything the agent ran, in all 21 sessions. Of 526 gh pr checks calls, 154 used --watch, and after the watch the agent asked two to four more times "to be sure".
The docs: Claude Code "treats them as context, not enforced configuration. To block an action regardless of what Claude decides, use a PreToolUse hook instead"(opens in new tab). The same page says there is "no guarantee of strict compliance."
The built-in prompt has the same problem. One bug report quotes it, "If waiting for a background task you started with run_in_background, you will be notified when it completes — do not poll", and the reporter writes that the model is "clearly not adhering to it"(opens in new tab).
What did the agent do after loop-guard refused its polling?
It found another way to spend a turn. In clipwright, loop-guard refused sleep followed by a reader such as tail or cat, and wait loops built on sleep. A bare pause was still allowed. The refusal fired twice in 11 sessions. Then one session, eda1, which ran 8 to 10 September, sent 273 echo waiting turns in four bursts. The largest was 180 turns between 21:19 and 21:38 on 9 September: 18 minutes 39 seconds, one about every six seconds. The other 93 came in three later bursts, 71 and 19 on 10 September and 3 on 9 September. At one point the agent announced, in words, that it was not polling.
None of those turns made the wait shorter. Each was a model turn spent printing a word.
Reports from outside describe the same shape. One traces it step by step: "Foreground sleep is blocked by the Bash tool, so the only wait the agent reaches for is a backgrounded one"(opens in new tab), which produced 42 backgrounded sleeps in one 33-minute run. Another describes a run that emitted about 500 no-op poll commands(opens in new tab) and notes that prompt-level guidance "does not mitigate this": the model quoted its own "wait, don't probe" instruction and spiraled anyway.
On 15 September we ported the hooks into our main project, and loop-guard carries four rules about waiting. Two were added after the 273 turns. One refuses an empty turn. The other refuses gh pr checks run in the foreground without --watch, which answers the gh polling, not the echo turns. Here is the empty-turn check, from .claude/hooks/loop-guard.mjs:
1export const NOOPS = new Set(["echo", "printf", "true", ":"]);
2
3export function isNoopWait(command) {
4 if (/[$`><]/.test(command)) return false;
5 const segs = segments(command);
6 if (!segs.length) return false;
7 if (!segs.every((s) => NOOPS.has(firstWord(s)))) return false;
8 const words = command.trim().split(/\s+/).filter(Boolean);
9 return words.length <= 6;
10}
A command is an empty turn only if every segment is a no-op, nothing is substituted or redirected, and it is short. The test file pins the edge: echo "=== x ==="; git status passes, because the second segment does real work.
The hook is written in Russian, so this is a translation of its output for the noop-wait rule, unabridged:
1BLOCKED (noop-wait): an empty turn (`echo waiting`, `true`, `sleep`): it does
2nothing, and it does not speed up the wait: 273 such turns in one session.
3
4Instead: nothing. The background command will send a task notification by
5itself, a subagent will send a message. If you are waiting for something
6external (CI, a deploy): `gh ... --watch` or Monitor, through run_in_background.
The message names sleep, though NOOPS leaves it out. A lone pause is allowed, and the comment in the source says so on purpose. The intent is that a refusal tells the agent the permitted move along with the reason. We have not measured whether that changes behaviour in our project.
What does cms-token.sh do about 546 secret reads?
It makes the repeated work unnecessary, and the cause was mechanical. Across our main project's sessions the pass counted 546 secret reads and 108 CMS logins. Shell variables do not survive between tool calls, so almost every command that touched the CMS rewrote the same four-line preamble: load the environment, read two secrets, log in, parse the token.
scripts/cms-token.sh replaces the preamble. It logs in once, caches the session in a file, and prints only the file's path. The script's header says why:
1# It prints the PATH of a 0600 curl config file, never the token. Pass it to
2# curl with -K. A transcript is a file on disk that outlives the session, so a
3# token echoed into one is a token leaked; the path is enough for every caller.
A caller writes curl -K "$(scripts/cms-token.sh da)" <url>. The cache expires at 100 minutes against Payload's default token lifetime of 7,200 seconds, which leaves 20 minutes of headroom. When we added it, a second call reused the cached session without a new login, and an authenticated request through the printed path returned 200. We have not recounted secret reads since.
What held, and what did not?
In clipwright, four of eight fixes held, and one was routed around. The table in its reflection, headed as seven fixes from 27 August, lists eight:
- Held (four): the review-round ceiling, with a hole in its accounting;
edit-channel-guard;doctor.sh, almost mute; andsession-handoff. - Held in part (one): stopping idle subagents. One session ran 38 stops for 76 subagents, and five sessions ran none.
- Never fired (one): a comment-budget hook, with 0 firings in 11 sessions, so no evidence either way.
- Routed around (one):
loop-guard, as above. - Did not take (one): a quiet-mode setting for a tool notice, which still appeared 4,026 times in those 11 sessions.
Two of the held ones have numbers. edit-channel-guard, which refuses repository edits made outside Edit and Write, fired twice in 11 sessions. Before it, heredoc and sed -i edits had caused two thirds of the blocks from a separate command guard, and afterwards they did not. The session-handoff hook, which prints branch, status and recent commits at session start, brought reorientation down to one or two turns. Those are clipwright results from before the port. The same hooks have run in our main project only since 15 September.
The subagent finding is the open one. Thirty-nine of 51 subagents in our main project were stopped by the owner at the end of a session, and 54 results arrived truncated, over 16,000 characters. The fix is a sentence: ask for 40 lines or fewer plus a file path holding the detail. Whether a hook could sweep idle agents at session end is a question in our notes, not a result. This is the fix most likely to fail, and we have not recounted our main project since the hooks landed.
How do you count your own in ten minutes?
Count the first word of every command your agent ran, and read the top of the list. Claude Code keeps transcripts as JSONL under ~/.claude/projects/, one folder per project. From your project's folder:
1jq -r 'select(.type=="assistant") | .message.content[]?
2 | select(.type=="tool_use" and .name=="Bash")
3 | .input.command | split(" ")[0]' *.jsonl | sort | uniq -c | sort -rn | head
First words are crude, but they surface the shape fast. Look for echo, sleep, true, or one command repeated across consecutive turns. Pick the one that repeats most inside a single session and write the guard for that behaviour, with the permitted move in its block message. Then recount in two weeks, and keep the transcripts so you can.
FAQ
Why does an agent ignore a rule in CLAUDE.md?
Claude Code delivers CLAUDE.md as a user message after the system prompt and treats it as context, not enforced configuration. Anthropic's memory documentation says there is no guarantee of strict compliance and points to hooks for anything that must hold.
Is blocking sleep enough to stop an agent polling?
Not in our clipwright session. The guard refused sleep followed by a reader, and wait loops built on sleep. The agent then sent 273 echo waiting turns. A guard has to cover the behaviour, which here includes an empty turn and a foreground poll, and say what to do instead.
What should a block message from a hook say?
Ours says three things: which rule fired, why it exists, and the permitted alternative. That is a design choice, and we have not measured whether it changes what the agent does next.
The counts cover 27 August to 15 September 2026 and come from a reflection pass over our own transcripts. The code excerpts are verbatim from our project, and the hook's block message is translated from Russian. This builds on why chained agent steps fail, why an agent's "done" needs a gate it cannot edit, and why one reviewer is not enough.
Which of your agent's rules does it actually follow?
Until you count, you find out one wasted turn at a time. No product of ours fixes that. We counted ours and some repeats are still unfixed. Choosing which rules deserve a guard? Tell us what you see.
About the Author
Dimantika
Co-founder of Dimantika. Builds Clipwright and ViralFaceless with coding agents. Previously ran GlockSoft with a partner for about 15 years. Writes about products and finding customers.
View all postsMore articles you might like.

The Agent Said "Done." It Wasn't.
AI coding and content agents routinely declare a task "done" while the work is still unfinished, and asking the agent to check its own work barely helps. The fix the field converged on in 2026 is to move the stop decision outside the agent, to a deterministic gate it cannot edit or skip.

STORM Without Retrieval Is Just Five Hallucinations in a Trench Coat
A viral thread promises Stanford's STORM research method in four prompts with no setup. It quietly deletes the one thing that makes STORM research instead of confident guessing — retrieval. Here's why grounding is the whole point, and how we keep it.

Failed Is Not a Terminal Status: A Contract for Agents
A run can fail on our own timeout while the paid work can still finish. Why an agent should not treat failed as a terminal status without a reason it can read.