ADR-089: Guardrail boundary completeness — reasoning, system prompt, vision
- Status
-
Accepted
- Date
-
2026-07-18
- Authors
-
Netresearch DTT GmbH
Context
ADR-085/087 guard the model's answer (output) and the user turns (input). An adversarial audit of the merged layer found the secret-redaction guardrail — a defense against a model echoing a secret it was given, or a user pasting one — missed three reachable boundaries, so the same secret masked in one place leaked in another:
- the model's reasoning / ``thinking`` block (surfaced in the playground glass-box, ADR-040) — screened on the answer, not the reasoning;
- the system prompt — added after input screening, so it reached the
provider unscreened on every path except
complete;For Configuration - the vision text prompt —
visionforwarded its text items to the provider without screening, on both sides.()
Decision
Cover every boundary the same secret can cross.
- Reasoning.
Guardrailgains an optionalResult redacted;Thinking Secretmasks both the content and theRedaction Guardrail thinkingblock, andGuardrailrebuilds the response with the redacted reasoning (null = leave as-is). Tool-call arguments are deliberately not redacted — they are functional parameters the tool consumes, and masking them would break the call.Middleware - System prompt. The six
applycall sites inSystem Prompt () Llmroute throughService Manager apply, which screens the final assembled message list — so the prepended system turn is screened too. Re-screening the already-screened user turns is an idempotent no-op.And Screen System Prompt () - Vision.
visionscreens the text items of its() Visionbefore dispatch, matching the chat/tool paths.Content
Consequences
- A secret is now masked (or the prompt blocked) across answer, reasoning, user turn, system prompt and vision text — the "enumerate all boundaries" completion of ADR-085/087.
embedremains unscreened: its input is embedding text, not a chat prompt; screening/redacting it would corrupt the vector semantics. Documented as out of scope.() - Redacting reasoning depends on the provider populating
thinking(Ollamamessage., ADR-016); providers that fold reasoning into the content are covered by the content redaction already.thinking - On the streaming paths the system prompt is screened inside the lazy opener
(that is where
applyruns), so a REDACT is applied before the provider send, but a DENY / REQUIRE_APPROVAL triggered only by system-prompt content would throw on first drain rather than at call time — unlike the eager user-turn screening. Latent: the shippedSystem Prompt Secretonly REDACTs, which is correct here; a future DENY-returning input guardrail would need the effective system prompt hoisted and screened before the opener.Redaction Input Guardrail