ADR-032: Specialized Usage Tracking and Pricing Catalog
- Status
-
Accepted
- Date
-
2026-06-10
- Authors
-
Netresearch DTT GmbH
Context
The chat/embedding path records complete usage rows: the middleware
pipeline (ADR-026: Provider Middleware Pipeline) tracks tokens and derives a cost from the
admin-curated tx_ pricing via Model::.
The specialised services bypass that pipeline by design — but they
recorded almost nothing. The image services passed metric keys
(size, quality, count) that
Usage does not map, so only
request_ landed in tx_: no cost, no
tokens, no images_, no model_. TTS recorded
characters but no cost; Whisper recorded nothing but the request.
Consequently the Analytics module, the MonthlyCost widget and
BudgetService systematically excluded all image and speech spend —
defeating the requirement that nr_llm can monitor total AI spend.
Two structural problems compounded this:
- the specialised services have no access to model pricing (their
models —
gpt-,image- 2 tts-,1 whisper-— usually have no1 tx_row), andnrllm_ model gpt-responses carry aimage-* usagetoken object (DALL·E responses do not), which was discarded.
Decision
- Real units in the callers. The services pass the metric keys the
tracker actually maps:
images(→images_),generated characters,audio(→Seconds audio_, from theseconds_ used verbose_Whisper duration), token keys when the response reports them, and the model identifier asjson model(→Id model_). Provider strings drop the ad-hocid provider:suffixes (model dall-→ providere: dall- e- 3 dall-+e model_).id - Token usage parsing.
Dallparses theEImage Service usageobject ofgpt-responses (image-* input_,tokens output_,tokens total_,tokens input_) so token aggregates include image calls; DALL·E responses withouttokens_ details usagegracefully omit token metrics. - Static price catalog with a DB override.
Specializedencodes the published OpenAI list prices (each constant documents source URL and verification date): gpt-image-* token prices and per-image fallback estimates, DALL·E per-image prices by quality/size,Pricing Open Ai Price Catalog tts-/1 tts-per 1M characters,1- hd whisper-per minute.1 Specialized(injected intoCost Calculator Abstract) resolves in order: admin-curatedSpecialized Service tx_row matching the model identifier (reusingnrllm_ model Model::, so negotiated prices win) → catalog token prices → catalog per-image price →estimate Cost () 0.. Unknown models never get a guessed cost — a zero cost signals "no price data" instead of fabricating numbers.0 - No double counting.
Llmno longer repeats the token count on its translation row (the pipeline already records tokens and cost on the underlying chat row); it keeps the translation-level request/characters view.Translator Whisperloses its secondTranscription Service:: translate To English () trackcall — the dispatch path records the request exactly once.Usage ()
Consequences
- ● Image, TTS and Whisper spend appears in the Analytics module, the MonthlyCost widget and BudgetService aggregates — total spend monitoring covers all service types.
- ● Costs follow published list prices and can be overridden per model
by creating a
tx_row with token pricing.nrllm_ model - ◑ The catalog requires manual maintenance when OpenAI changes list prices; constants carry source URLs and verification dates to make the review mechanical.
- ◑ Analytics grouped by
service_now showsprovider dall-/e fal/tts/whisperinstead of suffixed variants (dall-); historic rows keep their old strings, the model dimension moved toe: dall- e- 3 model_.id - ◑ FAL calls record images but cost
0.— FAL publishes no static list prices for its hosted models.0