---
title: "ADR-211: Typed decisions from an exchangeable backend"
manual: "TYPO3 LLM Extension"
version: "0.38"
permalink: "https://docs.typo3.org/permalink/netresearch/nr-llm:adr-211@0.38"
source: "Adr/Adr211TypedDecisionsFromAnExchangeableBackend.rst"
rendered: "2026-09-30T10:31:37+00:00"
---

# ADR-211: Typed decisions from an exchangeable backend {#adr-211}

-   *Status:* Accepted
-   *Date:* 2026-09-28
-   *Amends:*

    [ADR-013](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-013@0.38) (prices stored as integers),
    [ADR-060](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-060@0.38) (the LLM judge grader),
    [ADR-082](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-082@0.38) (what the structured methods return),
    [ADR-128](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-128@0.38) (what their callers consume),
    [ADR-129](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-129@0.38) (the judge as a structured consumer)

-   *Authors:* Netresearch DTT GmbH

## Context {#context}

Consumers of nr-llm increasingly need a *judgement* rather than a text: is
this passage relevant to the question, does this script keep to its source,
which of these five categories fits, how complete is this answer on a
four-level rubric. Today there is exactly one place that asks a model for
such a judgement, `LlmJudgeGrader` ([ADR-060](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-060@0.38)), and it is
an offline evaluation grader: it asks the default chat configuration for a
`{"score", "reason"}` verdict, and folds a failed call into a failed grade
with score 0. ADR-060 names "selecting a dedicated judge model" as an open
follow-up.

A consumer that wants a judgement at request time has no contract for it. It
can call `completeStructured()` with its own schema, which works, but every
consumer then invents its own question format, its own reading of a
self-reported "confidence", and its own answer to "what does a failed call
mean".

Models now exist that make such decisions natively. In September 2026
TypeSafe published Jev (`POST /v1/systemone`), a model that generates no
text: it takes a `state` and a map of typed questions — `noul` (yes/no,
answered as the probability of yes), `choice` (one option out of up to 255,
with a probability per option and a `confidence`) and `score` (2 to 10
ordered levels, a probability-weighted value, a probability per level and a
`confidence`) — and bills input tokens only. Its documentation is explicit
that a `noul` answer carries no `confidence`, that aliases such as
`jev-latest` move without notice, that English is the primary training
language, and that adversarial content in the state can move the answer.
Freely available natural-language-inference models answer the same three
question shapes locally as zero-shot classification — for example
`MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli-2mil7` (MIT, trained
on German among other languages) — with probabilities that are the model's
distribution, not a calibration.

The existing code decides a good part of the shape:

-   Every model an administrator can choose is a `Provider` record, a
    `Model` record with capabilities, and an `LlmConfiguration` that a
    consumer names per use case. Budget per configuration, usage and pricing
    per model, fallback chains, the circuit breaker on provider exceptions and
    the provider's trust zone all hang on these records.
-   `ProviderInterface` requires chat, completion and embedding methods;
    a provider that lacks one throws `UnsupportedFeatureException` there
    (Claude and Groq for embeddings), and optional abilities are separate
    contracts (streaming, tools, vision, documents).
-   `Model` stores prices as integer cents per million tokens. Jev's
    0.042 USD is 4.2 cents and cannot be stored.
-   `UsageMiddleware` records a typed response per call on the telemetry
    row ([ADR-174](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-174@0.38)), with `NULL` where nothing was measured.
-   `GuardrailInterface::checkOutput()` receives only the
    `CompletionResponse`, and `InputGuardrailInterface::checkInput()`
    one string. Neither sees the task or the evidence an answer should be
    judged against, so a judgement "does this answer the question from these
    sources" cannot be a guardrail without a new contract.
-   `completeStructuredForConfiguration()` runs a schema-bound call against a
    named configuration, but returns only the decoded array: the model that
    answered and the tokens of a repair round-trip are lost.

## Decision {#decision}

**A decision is a model operation like chat or embeddings. A model that
decides natively is a** `Model` **record with the capability**
`decision` **behind a provider adapter that implements**
`DecisionCapableInterface`**; any chat model can answer the same
questions through structured output. A new** `@api` **service,**
`DecisionServiceInterface`**, evaluates a subject against a named,
versioned profile of typed questions on a configuration and returns typed
answers.** The feature is called *decision*, not *judge* and not after a
vendor: classification, selection and rubric scoring are one operation.

1.  **Neutral question types.** `YesNoQuestion`, `ChoiceQuestion`
    and `ScoreQuestion`, with `DecisionSubject` and
    `DecisionAnswer`, are domain value objects, because a provider
    contract reads them. Their constructors enforce the limits a backend
    would otherwise reject after a round-trip: a choice has 2 to 255 unique
    options, a score 2 to 10 levels, a key is a lower-case identifier. A
    choice takes its option names as a list and their descriptions as a
    separate map: one map of name to description cannot tell
    `['red', 'green']` from options named `"0"` and `"1"`, which PHP
    stores under the same integer keys. Vendor vocabulary (`noul`) stays
    inside its adapter.
1.  **Profiles are declared by the consumer, the model by the operator.** A
    consumer extension implements `DecisionProfileProviderInterface`
    (tag `nr_llm.decision_profile`, the discovery pattern of
    [ADR-056](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-056@0.38) and [ADR-060](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-060@0.38)). A
    `DecisionProfile` carries an identifier, an integer version, its
    questions, the subject fields it requires (`task`, `candidate`,
    `evidence`) and the **data class** of its subject
    (`ToolDataClass`, default editor content). A request names a
    configuration, or the service uses the one in the extension setting
    `decision.configuration`; the default chat configuration is never
    used. A consumer therefore picks a decision model per use case exactly
    as it picks any other model, and never holds a key.
1.  **Data policy through the trust zone.** Before any request, the service
    checks the profile's data class against the least trusted zone the call
    can reach — the provider of the model that will serve it and every
    provider one fallback hop away, a criteria-mode fallback counted as
    external-global (`TrustZoneResolver`, the ceiling
    [ADR-094](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-094@0.38) introduced for tools). The model is resolved
    once and that routing decision is handed to the call, so the model
    checked is the model that serves; the input-context gate reads the same
    decision. A profile whose subject is
    internal configuration is refused on an external-global provider,
    whichever model that is. This replaces a list of permitted backend names,
    which said nothing about where a backend sends its data.
1.  **Two paths, chosen by the model.** The service resolves the
    configuration's model for `ProviderOperation::Decision`. Where that
    yields no decision model — a fixed chat model, or criteria that match no
    decision model — the structured path runs as the chat call it is: it
    resolves for `ProviderOperation::Chat` (a model that declares `chat`,
    or one that declares no capabilities at all) and is
    recorded under that operation, so telemetry and usage name what
    actually ran.

    -   *Native* — the model declares `decision`; its adapter must implement
        `DecisionCapableInterface`, and one that does not is refused as a
        model that cannot decide rather than asked through the other path:
        `LlmServiceManager::decideForConfiguration()` screens every subject
        field through the input guardrails, enters the middleware pipeline
        with the configuration (budget per configuration and user, fallback,
        circuit breaker, telemetry, usage), and calls the
        adapter. The adapter returns a typed `DecisionResponse`; the
        manager prices it with the model that actually served, fallback
        included. Two adapters ship: `typesafe` (Jev, pinned default model
        `jev-1.13.0`, because thresholds tuned on one version must not move
        under a caller) and `decision_sidecar`, a local zero-shot NLI service
        in `Build/decision` that needs no key and keeps the subject on the
        host — for tests, local development and comparison.
    -   *Structured* — any other model, typically a chat model: a schema-bound
        call through `completeStructuredForConfiguration()` asking every
        question for a hard label (`yes`/`no`, one option, one level index),
        because a probability a chat model writes into its text is not a
        measured one.

    The structured path needs what the structured call so far threw away: the
    model that actually answered and the tokens of every attempt, the
    rejected first answer of a repair included. So `completeStructured()`
    and `completeStructuredForConfiguration()` **return a**
    `StructuredCompletionResponse` — the validated `data`, the
    accepted `CompletionResponse`, the summed `UsageStatistics` and
    the number of attempts — instead of the bare array. This is a breaking
    change for every caller, taken instead of a second method beside each of
    the two.
1.  **Prices are decimal.** `Model` keeps its unit, cents per million
    tokens, and stores it with two decimals — the precision TYPO3's backend
    form keeps for a decimal field — so a price of 4.2 cents is a price and
    not 4 or 0; the smallest price stored is 0.01 cents per million tokens.
    This is a breaking change of `Model`'s price getters from `int` to
    `float`, and of the criteria cap on the input price with them.
1.  **The result says what was measured.** `DecisionResult` carries, per
    question, the answer (the probability of yes, the chosen option, the score
    value), the per-option or per-level probabilities and the `confidence`
    **as the model reported them, empty or null where it reported none**. Its
    `ProbabilityKind` names which kind a caller holds: `Calibrated`
    (a vendor's calibrated distribution, TypeSafe), `Distribution` (the
    model's own probabilities, uncalibrated — the NLI sidecar) or `None`
    (hard labels, the structured path). The result names the configuration,
    the provider and the model that answered; token counts and cost are
    `null` when nothing reported them (a measured zero and an absent
    measurement stay distinct, constitution principle VI).
1.  **A failure is an exception, never a result.** An unknown profile or
    configuration, a missing subject field, a data class the provider may not
    receive, a model that can neither decide nor chat, a request the provider
    rejects as invalid, a transport failure and an answer that does not match
    the questions asked — wrong key, type, option, level range or probability
    keys — all throw `DecisionException` with a named code. Budget,
    guardrail and input-context trust-zone denials keep their own types, as
    policy a caller may handle. No code path turns "the model could not be
    asked" into an answer. A fallback configuration whose model cannot serve
    the operation at all is skipped rather than reported in place of the
    primary's failure — for every operation, since the same holds for a
    sibling that lacks embeddings or vision.
1.  **No verdict, no enforcement.** The service returns answers. Thresholds,
    what follows from an answer (warn, block, retry, ask a human) and every
    permission check stay in the caller's code.
1.  **Criteria come from the profile, never from the subject.** The
    instructions, options and levels a model receives are the profile's, so a
    document under judgement cannot replace the rubric it is sent with. It can
    still try to sway the answer — TypeSafe documents that, and a chat model
    reads the subject in the same prompt as the questions. This narrows the
    manipulation risk and does not close it, which is the reason for point 8.
1.  **An operation and a capability.** `ProviderOperation::Decision`
    labels the native call and maps to `ModelCapability::DECISION`, which
    the operation map enforces for criteria-mode selection. Model discovery
    writes the capability for the two decision adapters, and the backend
    module's model and configuration tests send a decision probe to a
    decision model instead of a completion it cannot answer. A decision model
    answers no chat call, so it must not serve generic `chat()` calls: the
    setup wizard never makes such a model the default model, and the
    operator documentation says not to make a decision configuration the
    default configuration. A chat call that still reaches one fails with
    `UnsupportedFeatureException` rather than an answer.
1.  **The decision grader replaces the LLM judge.** `DecisionGrader`
    (grader identifier `decision`) grades a golden prompt through the
    built-in profile `nr_llm.task_fulfilment` on the configured decision
    configuration, and `LlmJudgeGrader` (`llm_judge`) is removed:
    keeping both would leave two judges with two failure semantics, one of
    them bound to the default chat configuration. The grader, not the
    service, folds a failure into a failed grade, because in an offline run
    one bad call must not abort the set. What would fail every grade — no
    configuration, a profile that cannot be used, a model that cannot answer,
    a trust zone that may not receive the task — is checked before the run
    spends its first completion (`assertAvailable()`), and refuses the run
    instead. The grader judges the response
    against the task and the system prompt the call ran with, so an ignored
    instruction counts against it.

    A run is stored and compared under the grader its gradings report, and
    the decision grader reports its yardstick,
    `decision:<provider>:<model>:v<profile version>`. Runs on two providers
    or two model versions are separate series and never each other's
    regression baseline. A run whose gradings disagree (a decision failed for
    some prompts) is stored under `decision` and never compared; a failed
    decision reports `decision:failed`, so no clean run shares either
    identifier. With `--fail-on-regression` both are a failure: a gate that
    stays green while the judge is down is the outage-read-as-pass this
    record rules out.

## Consequences {#consequences}

-   A consumer asks for a judgement through one typed contract and picks the
    model per use case like any other; an operator changes the model — TypeSafe,
    a local NLI sidecar, any chat model — without a code change in the
    consumer, and the trust zone decides where a subject may go.
-   A decision call is budgeted, priced, observed and failed over like every
    other model call, through the same records.
-   TypeSafe is a new external data recipient. It is reached only through a
    configuration an operator creates, receives only screened subject fields,
    and a profile's data class can keep a subject away from it. Region,
    retention and contract questions for confidential data are the operator's
    to settle; nothing here answers them.
-   Breaking: `completeStructured*()` returns a
    `StructuredCompletionResponse`; a caller reads `->data`.
-   Breaking: `Model` prices are `float` cents per million tokens.
-   Breaking: `--grader llm_judge` and `LlmJudgeGrader` are gone. Stored
    `llm_judge` results stay in `tx_nrllm_eval_result` but are no
    regression baseline for `decision` runs.
-   `LlmServiceManagerInterface` gains `decideForConfiguration()`; an
    implementation outside nr-llm must add it.
-   `ProviderOperation::Decision` and `ModelCapability::DECISION` are new
    cases. A consumer that matches either enum exhaustively must add them.
-   The structured path's result carries the tokens of every attempt and a cost
    only where the provider reported one; the priced cost of a chat call is in
    its usage record. Surfacing it on the response is a change to the chat
    pipeline, not to this service.

Follow-ups, each with the reason it is not here:

-   **A profile-to-configuration assignment in the backend.** Today a consumer
    passes a configuration or the extension setting names one for all
    profiles. An assignment per profile needs its own record and module view.
-   **Guardrail, cache and streaming integration.** A decision needs task and
    evidence that neither guardrail contract carries; a cached decision would
    need profile version, model and access context in its key; a check after a
    streamed response prevents nothing.
-   **Enforcement modes, cascades, routing signals, agent checkpoints and a
    reranker on a decision model.** They consume this service, and each needs
    measured error rates per profile first — which `--grader decision` on
    TypeSafe, the NLI sidecar and a chat model now produces, German content
    included.

## Alternatives considered {#alternatives-considered}

**Decisions as a specialized service configured in the extension settings**
(the first draft of this record: one backend per installation, the key in
`decision.typesafe.*`). Rejected: without a `Model` record a consumer
cannot pick a decision model per use case, an installation cannot run two,
there is no budget per configuration, the circuit breaker never trips on a
specialized-service exception, and the only data policy was a list of backend
names. The review of that draft found each of these as a separate defect.

**A** `ProviderInterface` **method for decisions.** Rejected:
`ProviderInterface` is an extension point that gains no abstract member
within a major version ([ADR-127](https://docs.typo3.org/permalink/netresearch/nr-llm:adr-127@0.38)); an opt-in contract beside
streaming, tools and vision is the established shape.

**A** `ModelType::JUDGE` **or a** `supportsJudge` **flag.** Rejected: a
generative model can judge, and a decision model cannot generate. The
capability states what the model can answer; the role is the profile's.

**A boolean** `calibrated` **flag.** Rejected: it cannot tell an NLI
model's own probabilities from a calibrated vendor distribution, and calling
the former calibrated would mislead every threshold built on it.

**A status field instead of exceptions** (evaluated / undeterminable /
unavailable / invalid). Rejected: every consumer would have to check it, and
one that did not would read an outage as an empty set of answers.

**Reuse or keep** `LlmJudgeGrader`**.** Rejected: its failure mode
(score 0) is wrong at request time, it is bound to the default chat
configuration, and beside the decision grader it would be a second judge for
one purpose.

**Follow** `jev-latest`**.** Rejected as the default: TypeSafe itself
advises pinning a version once thresholds are tuned. An operator can still
name the alias as the model id.
