The hardest agent failures are not always the ones that crash.
The harder class is when a run looks complete, the timeline looks plausible, or the final message claims success while the execution record says something else.
That distinction matters in production. If a transcript says a tool failed, you need to know whether a real tool attempt failed or whether the interface synthesized a placeholder. If a model emitted a tool call, you need to know whether the runtime actually dispatched it. If a workflow resumed, you need to know whether it resumed the same durable identity. If a planner marks a task done, you need receipts that prove the required action happened.
Here are four concrete failure modes worth adding to an agent reliability checklist.
1. A tool-call card does not always mean a tool ran
In one public Zed issue, an unknown ACP tool-call update could create a persistent failed tool-call card. To a developer reading the timeline, that card can look like a real failed attempt.
The debugging trap is treating an interpreted interface artifact as causal evidence.
- Does every visible tool-call card map to a raw protocol event?
- Does it map to a real dispatch attempt?
- Was it created as a synthetic placeholder after an unknown ID update?
- Does the interface distinguish raw events, inferred state, and real execution?
The debugger should preserve provenance. A synthetic placeholder can be useful, but it should not masquerade as the thing that actually happened.
Source: Zed issue #61667
2. A valid tool call can be received and still never execute
In one public LibreChat issue, valid tool calls from an OpenAI-compatible endpoint were received, but the tool execution step and continuation turn did not happen. The run ended with an empty response instead of an obvious runtime error.
That creates a nasty investigation path: the model did ask for the tool, but the application did not complete the tool-call state machine.
- Was the tool call received from the model?
- Was it validated against the registered tool schema?
- Was it dispatched to the tool implementation?
- Was the tool result appended back into the conversation?
- Did the system request the expected continuation turn?
- Did the final response reflect the tool result?
“Tool call received” and “tool executed” need to be separate states. Collapsing them hides the most important boundary in the run.
Source: LibreChat issue #14435
3. Resume can break when workflow identity changes
In public Mastra DurableAgent reports, a suspended workflow could resume with the wrong inner workflow identity. The system kept asking again because the resumed path was tied to a different snapshot key than the original suspended run.
From the outside, this can look like flaky memory or an unhelpful model. The deeper issue is identity handoff.
- Outer durable run ID and inner workflow run ID
- Snapshot key and resume payload
- Suspended step ID
- Parent and child workflow relationship
- The exact transition where identity changed
Durable agents need debugger views that show identity across suspend and resume boundaries. Otherwise teams compare symptoms instead of the state handoff that caused the loop.
Sources: Mastra issue #20213 and the reproduction repository
4. Final success is a claim, not proof
In one public Robonix Pilot issue, an embodied task could be marked complete from the planner model response even though the required action was never planned or executed.
That is the core lesson for agent debugging: a final success state is only a claim until execution receipts prove it.
- Which obligation did the user or workflow create?
- Which capability call should satisfy that obligation?
- Was that call planned and executed?
- Is there a durable receipt?
- Did a fresh environment observation confirm the result?
- Did the completion check compare the final claim with the required action?
This applies beyond embodied agents. Any agent that writes tickets, sends email, updates records, triggers workflows, or changes external systems needs receipts.
Source: Robonix issue #189
The practical checklist
Before accepting an agent run as successful, verify the run against execution evidence:
- Every visible tool event maps to raw protocol evidence.
- Every received tool call has a dispatch record or a clear skipped reason.
- Every dispatched tool has an output, error, timeout, or cancellation state.
- Every tool result led to the expected continuation path.
- Every resume step preserves the expected workflow identity.
- Every completion state is backed by receipts and current observations.
- Synthetic interface placeholders are labeled separately from real execution attempts.
Agent debugging is evidence reconstruction
The transcript is useful, but it is not enough. A model message, interface card, workflow status, or final answer can all be true in one layer of the system and misleading in another.
The debugger has to join those layers: model intent, runtime dispatch, tool result, durable state, side effect, observation, and completion logic.
The next time an agent run says it succeeded, ask a sharper question: What proves it?
Debug agent runs from evidence, not claims
Opswald is building debugging infrastructure for AI agent runs where transcripts are only one layer of the story.
Request early access →