---
title: "ADR-123: One catalogue of secret shapes for every masking path"
manual: "TYPO3 LLM Extension"
version: "0.35"
permalink: "https://docs.typo3.org/permalink/netresearch/nr-llm:adr-123@0.35"
source: "Adr/Adr123OneSecretShapeCatalogue.rst"
modified: "2026-09-16T22:09:16+00:00"
---

# ADR-123: One catalogue of secret shapes for every masking path

-   *Status:* Accepted
-   *Date:* 2026-07-29
-   *Authors:* Netresearch DTT GmbH

## Context

Three places in this extension masked secrets, and each knew a different subset
of what a secret looks like:

-   `RedactsSecretsTrait`, behind the response and prompt guardrails
    ([ADR-085](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-085@0.35) / [ADR-087](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-087@0.35)), knew modern OpenAI
    project keys, classic and fine-grained GitHub PATs, AWS and Google keys, Slack
    tokens and bare JWTs.
-   `ContentRedactor`, which decides what gets **written to the database** at
    privacy level REDACTED ([ADR-064](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-064@0.35)), knew only credential-bearing
    URLs, `Bearer` headers, a narrower `sk-` pattern and e-mail addresses.
-   `GetEnvTool` matched on variable **names** only.

Measured against twelve secret shapes, the guardrail masked eleven and the
privacy redactor seven fewer. The consequence was not theoretical: a secret
correctly stripped from a prompt on its way to a provider was still persisted in
cleartext, because the two paths disagreed about what a secret is.

`GetEnvTool`'s name-only rule had a second version of the same problem. The
tool is admin-only but **enabled by default**, and its output egresses to the
configured LLM provider. A variable whose name gives nothing away leaked its
value verbatim:

```text
GITHUB_PAT=ghp_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
STRIPE_LIVE=sk_live_<24 alphanumerics>
```

Neither name contains `PASS`, `KEY`, `SECRET` or `TOKEN`, so neither was
redacted.

## Decision

Move the shapes into one place, `Netresearch\NrLlm\Utility\SecretShapeRedactorTrait`,
next to the `ErrorMessageSanitizerTrait` it builds on, and have all three
consumers read from it.

### Two entry points, opposite failure modes

`preg_replace()` returns `null` when the regex engine gives up, and a bare
`(string)` cast turns that into `''`. On a redaction path, wiping the entire
content looks exactly like a successful, very thorough redaction. The two kinds
of caller need opposite handling, so the trait offers both explicitly:

-   `redactSecretShapes()` — **fails open**, keeping the text. For the
    guardrails: losing a model's whole response, or an outgoing prompt, because one
    pattern hit a backtrack limit is worse than missing that pattern.
-   `redactSecretShapesStrict()` — **fails closed**, returning null. For
    `GetEnvTool`: a value the redactor could not fully inspect is withheld
    rather than forwarded to a third party.

### `GetEnvTool` checks both name and value

The name rule is kept and the value rule is added, because each catches what the
other misses: a name rule catches an empty or unrecognised-format secret
(`DB_PASSWORD=hunter2`) that no shape pattern would match, and a value rule
catches a recognised secret under a neutral name. A test asserts that the three
neutral fixture names are **not** matched by the name pattern, so the value path
cannot silently stop being exercised if someone widens the name rule later.

The tool keeps its own, stricter URL-userinfo pattern, which masks the whole
`user:password@` rather than just the password. In a provider error message a
username is useful context; in a listing that egresses to a third party it is
half a credential. Tightening the shared trait instead would have silently
changed what every provider error message discloses.

### E-mail masking stays with the privacy redactor

An address is personal data, not a secret, and the guardrails must not begin
stripping addresses out of prompts and responses — removing one changes what the
text says. `ContentRedactor` masks them because it writes to storage; the
guardrails do not.

## Consequences

-   The privacy redactor now masks every shape the guardrails do, so the
    persist path can no longer be weaker than the egress path.
-   `GetEnvTool` no longer leaks secret-shaped values under harmless names. A
    connection-string variable still shows its host and path — the context the
    tool exists to provide.
-   New shapes (Stripe secret and publishable keys, SendGrid) were added while
    consolidating, so all three consumers gained them at once.
-   Adding the next shape is a one-line change in one file instead of three
    edits by someone who has to know all three places exist.
-   This remains best-effort. It recognises these shapes and nothing else, and does
    not weaken the rule that secrets belong in nr-vault, never in a prompt, a
    column or an environment variable.

## The catalogue lives in nr-vault

nr-vault carried the same knowledge for its plaintext scanner, and it had drifted
the other way — it knew Stripe, SendGrid, Twilio, Mailchimp and PayPal but not
OpenAI project keys or fine-grained PATs. The merged superset now lives upstream
in `NetresearchNrVaultSecretSecretPatternLibrary` (nr-vault ADR-031),
which is the right long-term home: one catalogue for every Netresearch extension
rather than one per extension.

This extension reads it. `SecretShapeRedactorTrait` iterates
`SecretPatternLibrary::all()` and `ErrorMessageSanitizerTrait` iterates
`SecretPatternLibrary::urlCredentials()`, so no secret regex is defined in
nr-llm any more. Both gained shapes in the process — Slack legacy tokens,
Mailchimp and SendGrid keys, Stripe publishable keys — without any change here.

Read through the library's STATIC methods, not the injectable
`SecretRedactorInterface`. These traits are used from objects the container
never builds: a provider exception, and
`ControllerBackendResponseErrorResponse`, which callers construct with
`new`. A trait that reached into the DI container to mask a string would be a
worse dependency than a static call to a pure pattern list.

Requires nr-vault ^0.13.0.

### Why the sanitiser stays a narrower subset

`ErrorMessageSanitizerTrait` deliberately applies only the two URL shapes,
not the whole catalogue. It runs on error messages, where the job is to strip the
credential a client library put into a request URL — not to scan arbitrary prose
for every vendor token shape. Callers that want the full catalogue use
`SecretShapeRedactorTrait`, and the guardrails do.
