Skip to content
Everruns Cloud is open in early access. Run agents without operating the platform.
IDoutput_truncation
CategorySafety
FeaturesNone
DependenciesNone
RiskLow

A model response that reaches its output token limit (finish reason length) can stop in the middle of a tool call. Everruns never runs such a call, and never runs it with empty ({}) arguments in place of the ones the model did not finish. This capability chooses what the turn does next. On OpenAI Responses and Bedrock it also applies when a finished call arrives with arguments that are not valid JSON.

Every agent gets the default policy, continue, without adding the capability. Add it only to change the policy or the retry limit.

None, this capability only configures the turn loop.

When a response loses tool calls this way:

  • Calls in the same response whose arguments arrived complete still run.
  • The cut-off calls do not run. Every provider drops them: Anthropic, OpenAI (Chat Completions and Responses), Gemini, and Bedrock.
  • The policy decides the turn’s next step:
PolicyWhat happens
continue (default)The model gets a short notice that its output hit the limit and the call was not run, and the turn runs another model call so it can retry with less output. After max_retries cut-off responses in a row, the turn fails with an error.
failThe turn ends with an error.
offThe turn carries on as if the model had not asked for the cut-off calls: it ends on the partial answer, with stop reason max_tokens.

The continue notice appears in the conversation the model sees, right after the cut-off response and any results of the calls that did run. It is not stored as a message and is not shown in the chat. A clean model response, or a new user message, resets the count of consecutive cut-off responses.

Responses that are refused or filtered (content_filter, refusal) are not retried: retrying does not fix them.

{
"capabilities": [
{
"ref": "output_truncation",
"config": { "policy": "continue", "max_retries": 2 }
}
]
}
FieldTypeDefaultDescription
policycontinue | fail | offcontinueWhat the turn does when a response loses tool calls
max_retriesinteger, 0 to 102Cut-off responses in a row continue retries before the turn fails

When the capability is set on both the harness and the agent, the agent’s config wins.

Each llm.generation event records how the response ended:

  • finish_reasons;
  • tool_calls_dropped;
  • truncation_gate, set to retried or failed when the policy acted.

The same value is exported as the everruns.llm.truncation_gate span attribute and counted in the everruns_llm_truncation_gate_total metric. See the event reference.