ADR-055: Embeddings join the configuration path; dimensions metadata
- Status
-
Accepted
- Date
-
2026-07-13
- Authors
-
Netresearch DTT GmbH
Context
The three-tier model (Provider → Model → Configuration,
ADR-001) reaches every chat-shaped capability:
complete, chat,
stream and
chat all resolve the adapter from a
DB-backed Llm (vault key + model + pricing) and run
through the middleware pipeline, so budgets are enforced and cost is
attributed per configuration.
Embeddings did not. Llm only accepted
Embedding with raw provider/model strings, resolved
against ExtensionConfiguration and a model-less transient
configuration. An embedding consumer that persists vectors (a search
index, semantic auto-linking — see the scope boundary in
ADR-050) therefore had to duplicate provider, model
and dimensionality into its own extension configuration, bypassing
per-configuration budgets and cost attribution entirely.
The dimensionality gap made this worse: no record anywhere stated how many dimensions a model's vectors have. A consumer validating a persisted vector index against the configured model had to run a live "calibration probe" — embed a throwaway string and count the floats — which costs a provider call and fails when the provider is unreachable.
Decision
Embeddings join the configuration path.
Llm mirrors
chat: it resolves the adapter via
get, runs through the middleware
pipeline with Provider and the budget
metadata from the options, and guards the embeddings feature the
same way embed does (Unsupported when the
provider lacks it). Per-call Embedding take precedence over
the configuration's stored defaults — an options model overrides
the configuration's model id. Caching mirrors embed: a positive
cache_ places a cache key on the call context (keyed on the
configuration identifier plus the effective model), so two
configurations pointing at different models never share cache entries.
The high-level feature service follows:
Embedding and
embed delegate to the manager and populate
be via the shared auto-populate wiring, exactly like the
existing embed/embed paths.
Model records carry dimensions metadata. tx_ gains
a dimensions column (integer, 0 = unknown, declared like
context_), surfaced in the TCA next to the other model
limits and on the Model entity as
get/set. It is descriptive metadata:
nothing in nr_llm enforces it at call time.
Consequences
- Embedding consumers select a backend-managed configuration instead of duplicating provider + model + dimensions into their own extension configuration; per-configuration budgets and cost attribution apply to embeddings like to every chat-shaped capability.
- A consumer can validate a persisted vector index against the
configured model by comparing its stored dimensionality with
get— no live calibration probe, no provider round-trip. A value ofLlm Model ()->get Dimensions () 0means "unknown"; consumers fall back to their previous behaviour then. LlmandService Manager Interface Embeddinggained methods — implementers outside this repo must add them.Service Interface - nr_llm's embedding capability remains stateless (ADR-050): the configuration path changes how the call is resolved and accounted, not what is persisted. Vector stores stay out of scope.