ADR-107: Agent-loop context-window management
- Status
-
Accepted (which model's window, and which paths bind one — see ADR-143)
- Date
-
2026-07-22
- Amended
-
2026-08-10 by ADR-143
- Authors
-
Netresearch DTT GmbH
Note
The mechanism below is unchanged. Two things around it are:
ADR-143 takes the window from the model that actually serves
a send rather than the configuration's stored relation — which is null on a
criteria-mode record, so those calls were sized against the unknown-model
fallback — and binds the generic Llm chat, completion and
streaming paths, which previously bound none.
Context
The tool loop (ADR-081 ff.) appends each assistant tool-call
turn and its tool_result messages to one transcript that is re-sent on every
iteration. Nothing bounded that transcript against the model's context window
(Model::), so a long agentic run — many tool calls, large
tool outputs — eventually overflowed the window and failed at the provider with
a raw 4xx that FailureClassifier maps to a generic
CLIENT_, indistinguishable from an auth or config error.
Decision
An optional Context collaborator on Tool bounds
the transcript before each provider send. Absent it (the lean test wiring) the
loop sends the full transcript exactly as before — every enforcement site is a
no-op.
- Turn-atomic pruning (the correctness crux). Pruning drops oldest WHOLE
turns — an assistant tool-call message together with ALL its tool_result
replies — so the tool-call/tool-result pairing the provider requires is never
broken. The head (the leading system run plus everything up to and including
the first user message) and the newest turn are never dropped, so the output
is a structurally valid, non-empty transcript that still carries the task and
the most recent context. A cheap post-fit pairing guard defers to the provider
rather than ever emit a known-orphaned request. Summarization was rejected: it
needs an extra provider call per prune (cost, latency, non-determinism through
Suspended); dropping is deterministic, cheap and provably safe.Run State - Over-counting estimator. No BPE tokenizer (no runtime dependency): a
content-class-aware
chars/— prose divides by 3.5, DENSE segments (tool JSON arguments, tool_result payloads, the tool-schema block) by 2.5, plus per-message and per-tool-call overhead. A calibration factor seeded above 1.0 scales the estimate and only ever grows toward the real prompt-token counts each provider call reports, so the estimate errs high throughout and never under-prunes into an overflow. The manager is stateful per run and self-resets on each loop's first send (a nullN last), so a single sharedUsage Toolnever carries one run's calibration into the next.Loop Service - Reserve + graceful failure. `
budget = context`, else the model output cap, else a proportional floor) and an unknown context length falls back to a conservative 8192. When even the pruned floor still exceeds the budget, no provider call is made: aLength - reserve - safety``, where the reserve is the response allocation (options ``max_ tokens Contextstops the loop on the newTruncated Exception Agent(non-retryable) — a legible terminus instead of a misclassified provider 4xx.Run Termination Reason:: CONTEXT_ TRUNCATED
Consequences
AgentgainedRun Termination Reason CONTEXT_(TRUNCATED is) — the documented minor-release growth path.Retryable () === false - Global default only, no per-configuration knob (YAGNI);
max/Tokens Per Day modelare the storage precedent if an override is ever needed.Selection Mode - Enforcement covers the three real provider-send sites in
run(the in-loop tool send, the no-tools plain completion, the cap-hit synthesis); every send goes through one of them, so no separate pre-loop assembly pass is needed. The plain-completion sites pass no tool schemas, so phantom schema bytes never inflate the estimate of a payload that will not carry them.Loop - Streaming is out of scope.
Streamingruns a separate single-shot pipeline that never callsDispatcher run; bounding a streamed request against the window (and reconciling its ownLoop () chars/heuristic) is a follow-up.4 - Observability: a pruning event is logged at info level; a dedicated inspector
Runfor the trace is a follow-up. TheStep CONTEXT_reason is the load-bearing operator signal and is on the result.TRUNCATED