Definition

An agent planning failure happens when the proposed steps cannot achieve the task as requested. The plan may start from a false assumption, leave out necessary work, or put dependent steps in the wrong order. Individual tool calls can still succeed. The problem is the path the agent chose through the task.

Simple example

A repository agent is asked to rename a configuration field used by an API and a background worker. Its plan covers the API parser, an example file, and API tests. Every planned edit succeeds, and the tests pass. But the worker still reads the old field, so it cannot load the new configuration.

The missing worker update was a gap in the plan. Repeating the same edits or rerunning the same tests would not fix it.

Why it matters

Successful tool results report success for individual calls. They do not tell you whether the agent chose all the actions the task required. If you only inspect tool errors, this kind of failure can look like a successful run until a later check or deployment exposes the omission.

Before execution, compare the plan with the task’s acceptance criteria. When a tool result reveals another dependency, revise the remaining steps. In an evaluation, check coverage and ordering alongside tool-call validity.

One important nuance

A tool error can be the first visible sign of a planning failure. If a plan tries to migrate data before creating the destination, a rejected call points back to the order of steps. If the order was sound but the service timed out, the plan may still be valid. Recovery must account for whether the operation may already have taken effect before retrying or replanning.