Direct answer
Trace each workflow from the user’s request to the final business outcome. Capture model calls, retrieval, tool actions, validation, human review, cost, latency, failures, and state changes under one run ID. Token counts and model latency describe components; they do not tell you whether the customer outcome completed correctly.
A green model call can sit inside a failed workflow
The model returns a valid response in 900 milliseconds. The CRM update then fails, the retry creates a duplicate, the notification never sends, and the user sees “complete.” Every model dashboard stays green.
AI systems are distributed systems. Their failure boundary includes queues, databases, vector search, third-party APIs, permissions, human review, and the deterministic code around the model.
Observability has to follow that chain.
Start with a durable run identity
Give every user request or scheduled job a run ID. Carry it through retrieval, model calls, tools, queues, approval steps, and state changes. Child spans may have their own IDs, but the business action needs one trace that survives asynchronous work.
Attach stable context: tenant, user or service identity, workflow version, model and prompt versions, environment, action type, and the expected outcome. Avoid copying sensitive content into every log. Use references, redaction, and access controls appropriate to the data.
Instrument quality and consequence
For model calls, record latency, usage, routing, cache behavior, output validation, refusal, and evaluation signals. For retrieval, record the query, filters, source versions, result IDs, scores, and whether the evidence was sufficient. For tools, record normalized arguments, authorization, attempts, provider responses, idempotency keys, and durable effects.
Then close the trace with the outcome: sent, posted, approved, reconciled, rejected, abandoned, or escalated. If a person corrected the result, record the reason in a structured form.
That final label is what turns telemetry into product learning.
Dashboards should expose operational questions
Leadership needs cost and successful outcomes by workflow and tenant. Product teams need correction, refusal, abandonment, and escalation rates. Engineers need latency percentiles, tool failures, queue depth, retry loops, and regressions by release. Risk teams need unauthorized attempts, high-impact actions, and evidence gaps.
One generic “AI dashboard” serves none of them well.
In CARA, Zenveus recorded query-level telemetry across topic, jurisdiction, persona, latency, knowledge-base sufficiency, external lookup, refusal category, and response status. The administration layer separated agent performance, user analytics, funnel reporting, and individual-query review.
In a separate sales automation, immediate alerts cut failure-detection time from 8–26 hours to under 30 seconds. The important improvement was not more logs. It was a path from failure to the person who could act.
Alerts need ownership and a playbook
Alert on conditions that demand action: repeated tool failure, quality regression, cost runaway, stalled queue, missing evidence, unauthorized access, or a high-consequence workflow that did not complete.
Every alert should identify the affected workflow, scope, first failure, current state, customer impact, and safe next step. Deduplicate repeated symptoms. Assign an owner and response expectation.
An alert without an owner is only a louder log.
The acceptance test
Choose one failed customer outcome from last week. Can the team reconstruct the request, evidence, model decision, tool actions, permissions, cost, failure point, retries, human involvement, and final state from one trace?
Then choose one successful outcome. Can the team prove it was correct, not merely completed?
If either investigation depends on searching five systems and asking the original developer, the product has telemetry but not AI observability.
Related Zenveus services: Elastic Infrastructure and Agentic AI Development