Decision service 

Typed decisions about a subject, by a profile the consumer declares, from the model of a configuration (ADR-211). The service answers yes/no, choice and score questions and decides nothing: thresholds, consequences and permission checks stay in the calling code.

A decision is a model operation like chat or embeddings. A model record that declares the capability decision answers natively, through its provider's decision API; any other chat-capable model answers the same questions through structured output. Which one answers is a matter of configuration, not of code.

interface DecisionServiceInterface
Fully qualified name
\Netresearch\NrLlm\Service\Decision\DecisionServiceInterface

Public DI service. Resolve it by interface.

evaluate(DecisionRequest $request): DecisionResult

Resolve the profile and check the subject against the fields it requires; resolve the configuration and its model; check the profile's data class against the trust zone of every provider the call can reach; ask the model; verify that there is exactly one answer of the right type and range per question.

param DecisionRequest $request

profile identifier, subject, optional configuration identifier, budget subject (beUserUid) and caller source

throws

DecisionException for every failure except the ones below — its code names the cause (see Failures); BudgetExceededException, GuardrailViolationException, GuardrailApprovalRequiredException and InputContextTrustZoneException keep their own types, because a caller may handle them as policy. A failure is never returned as an answer.

Returns

DecisionResult

assertAvailable(string $profile, ?string $configuration = null): void

Check what needs no subject — the profile, the configuration, a model that can answer it and the trust zone against the profile's data class — without asking anything. A caller about to spend many paid calls whose results only a decision can judge checks this first, as nrllm:eval:run --grader decision does.

throws

DecisionException with the code evaluate() would fail with

Declaring a profile 

A consumer extension implements ProfileDecisionProfileProviderInterface; autoconfiguration tags it nr_llm.decision_profile. A profile carries an identifier, a version, its questions, the subject fields it requires and the most sensitive data class its subject can hold.

EXT:my_ext/Classes/Decision/MyDecisionProfiles.php
use Netresearch\NrLlm\Domain\Enum\ToolDataClass;
use Netresearch\NrLlm\Domain\ValueObject\Decision\ChoiceQuestion;
use Netresearch\NrLlm\Domain\ValueObject\Decision\ScoreQuestion;
use Netresearch\NrLlm\Domain\ValueObject\Decision\SubjectField;
use Netresearch\NrLlm\Domain\ValueObject\Decision\YesNoQuestion;
use Netresearch\NrLlm\Service\Decision\Profile\DecisionProfile;
use Netresearch\NrLlm\Service\Decision\Profile\DecisionProfileProviderInterface;

final class MyDecisionProfiles implements DecisionProfileProviderInterface
{
    public function getDecisionProfiles(): array
    {
        return [
            new DecisionProfile(
                identifier: 'my_ext.rag_answer',
                version: 1,
                questions: [
                    new ChoiceQuestion(
                        'verdict',
                        'What should happen to `candidate`?',
                        ['publish', 'revise', 'discard'],
                        ['revise' => 'Right in substance, needs rework'],
                    ),
                    new YesNoQuestion(
                        'supported',
                        'Is every claim in `candidate` supported by `evidence`?',
                    ),
                    new ScoreQuestion(
                        'answers_task',
                        'How completely does `candidate` answer `task`?',
                        ['Not at all', 'Partly', 'Completely'],
                    ),
                ],
                requires: [
                    SubjectField::Task,
                    SubjectField::Candidate,
                    SubjectField::Evidence,
                ],
                dataClass: ToolDataClass::EDITOR_CONTENT,
            ),
        ];
    }
}
Copied!

The questions:

  • YesNoQuestion — optional yesMeans and noMeans criteria.
  • ChoiceQuestion — 2 to 255 unique option names as a list and the descriptions as a separate map keyed by option, so options named "0" and "1" are names like any other.
  • ScoreQuestion — 2 to 10 levels; the answer is a level value from 0 to the highest level index.

A question key is a lower-case identifier of at most 64 characters ([a-z][a-z0-9_]*): it becomes a key of the answer map and of every wire format.

Raise the version whenever a question, an option, a level or an instruction changes: a threshold tuned against one version does not carry over. Criteria come from the profile only — the subject is data and cannot replace the rubric it is sent with. It can still try to sway the answer: TypeSafe documents that adversarial content moves its judgement, and a chat model reads the subject in the same prompt as the questions. Treat a decision as information, never as a permission.

A provider whose profiles cannot be built — one throws, or declares an invalid profile — is set aside, and an identifier declared twice is withheld; the other providers' profiles keep working, and asking for a withheld one names the failure (INVALID_PROFILE).

The data class is checked before anything is sent: a profile that holds SOURCE_CODE cannot be asked on a configuration whose provider — or any provider in its fallback chain — sits in a trust zone that may not receive it (Effective policy).

Asking for a decision 

Evaluating a RAG answer before it is shown
use Netresearch\NrLlm\Domain\ValueObject\Decision\DecisionSubject;
use Netresearch\NrLlm\Service\Decision\DecisionRequest;

$result = $decisionService->evaluate(new DecisionRequest(
    profile: 'my_ext.rag_answer',
    subject: new DecisionSubject(
        task: $question,
        candidate: $answer,
        evidence: $passages,
    ),
    configuration: 'rag_judge',
    callerSourceExtension: 'my_ext',
    callerSourceOperation: 'rag_answer',
));

$supported = $result->answer('supported');
// Your threshold, your consequence. Calibrate it per profile and model.
if ($supported->value < 0.8) {
    // e.g. show the answer with a warning, or not at all
}
Copied!

The configuration is the one the request names, otherwise the one the extension setting decision.configuration names — never the default configuration, so a change of the generating model does not change the model that judges it. Without either, the call fails with NO_CONFIGURATION.

Without a beUserUid in the request the service asks the backend user context, so the budget of the logged-in editor applies as it does to a chat call.

Which model answers 

The service resolves the configuration's model for the operation decision:

  • A model that declares the capability decision answers natively, through LlmServiceManagerInterface::decideForConfiguration(). The subject fields are screened by the input guardrails, and the call runs the configuration's middleware pipeline — budget, fallback, circuit breaker, telemetry and usage — under the operation decision.
  • Any other model that can chat — or that declares no capabilities at all — answers through a schema-bound structured completion on the same configuration. Every question is asked for a hard label (yes/no, one option, one level index), because a probability a chat model writes into its answer is text, not a measured distribution. The call resolves and is recorded as the chat call it is — operation chat, not decision — and input and output guardrails, budget, fallback and usage apply as to any chat call.
  • A model that declares capabilities but neither decision nor chat (an embedding model) cannot answer: MODEL_CANNOT_DECIDE.

In criteria mode the operation decision itself narrows the choice to models that declare decision — with operation capability enforcement on, the default (routing.operationCapabilityEnforcement); in observe mode a better-ranked chat model can win and answers through structured output. Where none matches, the service resolves the configuration for chat and asks that model through structured output; add cap:decision to the criteria to make a configuration decision-only. Either way the model is resolved once, and the call is served by exactly the model the trust zone was checked against.

Reading the result 

class DecisionResult
Fully qualified name
\Netresearch\NrLlm\Service\Decision\DecisionResult

profile, profileVersion, configuration (the one that was asked), provider and model (as the provider reported them), probabilityKind, answers keyed by question key, and inputTokens / outputTokens / cost — null when the call did not report them, never 0 in their place.

cost is set on the native path, by the provider or from the price of the model that served; on the structured path only where the completion reported one. The usage record carries the priced cost in both cases.

answer(string $key): DecisionAnswer
throws

DecisionException (NO_SUCH_ANSWER) for a key the profile has no question for

class DecisionAnswer
Fully qualified name
\Netresearch\NrLlm\Service\Decision\DecisionAnswer

value: the probability of yes (yes/no) or the level value (score, possibly between two levels); choice: the chosen option (choice); probabilities and confidence: exactly as the model reported them, empty or null where it reported none. Nothing is derived.

enum ProbabilityKind
Fully qualified name
\Netresearch\NrLlm\Domain\ValueObject\Decision\ProbabilityKind

What the probabilities of a result are worth.

None
Hard labels only — a chat model asked through structured output.
Distribution
The model's own distribution, not calibrated — the local decision sidecar.
Calibrated
A distribution the provider calibrates — TypeSafe.
Model yes/no choice score probabilityKind
TypeSafe (typesafe) probability of yes, no confidence option, probability per option, confidence weighted value, probability per level, confidence Calibrated
Local sidecar (decision_sidecar) probability of yes option, probability per option weighted value, probability per level Distribution
Any chat model 1.0 or 0.0 option level index None

A TypeSafe confidence of 0.95 is a property of one answer's distribution, not a measured accuracy of your decisions. Measure thresholds per profile, model and language — nrllm:eval:run --grader decision runs a golden set on whichever configuration decision.configuration names (Graders).

Failures 

DecisionException code Meaning
UNKNOWN_PROFILE No profile with that identifier is declared.
INVALID_PROFILE A profile provider declares an invalid profile or an identifier twice.
MISSING_SUBJECT_FIELD The subject lacks a field the profile requires.
NO_CONFIGURATION Neither the request nor decision.configuration names one.
UNKNOWN_CONFIGURATION No active configuration with that identifier exists.
DATA_CLASS_NOT_PERMITTED The profile's data class may not reach a provider the call can reach.
MODEL_CANNOT_DECIDE The configuration resolves no model that can answer, or its model declares decision on a provider that cannot make any.
REJECTED The provider refused the request — any 4xx but 429: a wrong key, a subject or body over the limit, a question it cannot take. Asking again unchanged does not help.
INVALID_ANSWER The model answered, but not with one valid answer per question — a chat model included whose reply missed the schema after the repair round-trip.
FAILED Every other failure: an outage, a timeout, a rate limit, an exhausted fallback chain. The cause is the previous exception. An \Error — a defect in code — is not wrapped and propagates.
NO_SUCH_ANSWER DecisionResult::answer() was asked for a key the profile does not have.

Decision models 

TypeSafe
Adapter type typesafe: TypeSafe's System One API (POST /v1/systemone). Create a provider with the endpoint https://api.typesafe.ai/v1 and the API key as an nr-vault identifier, then a model with the capability decision. Model discovery offers the pinned version jev-1.13.0 — recommended, because the aliases jev-latest and jev-preview move without notice — and prices it at the published 0.042 USD per million input tokens, output free. The subject leaves the installation: settle region, retention and contract for confidential content before enabling it, and set the provider's trust zone accordingly.
Local decision sidecar
Adapter type decision_sidecar: a zero-shot natural-language-inference model served on the host (Build/decision/), by default the multilingual MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli-2mil7. It needs no key, carries no price and keeps the subject on the host — the model for tests, local development and a comparison with a hosted one. Its probabilities are the model's own distribution, and bare yes/no questions lean towards yes; state what yes means. As a local service it refuses a few requests the questions themselves allow — an empty subject, more than DECISION_MAX_QUESTIONS (32) questions, choice options that render to the same label — with a 422, which the service reports as REJECTED. Like Ollama it is reached through a private hostname, so add that host (decision when the service carries that name on the container network) to $GLOBALS['TYPO3_CONF_VARS']['HTTP']['allowed_hosts'] (Testing a connection). It speaks plain HTTP; anywhere beyond a private container network, put a TLS-terminating reverse proxy in front of it.
Another native decision model
Implement NetresearchNrLlmProviderContractDecisionCapableInterface — or extend NetresearchNrLlmProviderAbstractDecisionProvider, which refuses chat, completion and embeddings and reads answers strictly — register the adapter type, and declare decision on its model records.

The backend module's model and configuration tests send such a model one yes/no probe instead of a chat prompt.

Configuration 

Extension configuration key (nr_llm, category decision):

decision.configuration
Identifier of the configuration a request without one is asked on. Empty by default: such a request fails with NO_CONFIGURATION.

Testing 

NetresearchNrLlmTestingFakeDecisionService returns queued results in order, records every request and throws a set throwable once — the way to test a caller's handling of a failed decision. It checks nothing: not the profile, not the subject's fields, not the answers. A test of that contract belongs against the real service.