Skip to main content

What it does

For every project and eligible metric, Cekura categorizes failing calls continuously as they come in:
  1. When a call fails a metric, its root cause is classified into that metric’s failure-mode set — reusing an existing mode when a single fix would resolve it, or opening a new mode when nothing fits. A call with more than one distinct root cause lands in multiple modes.
  2. Each failure mode carries a running count and its example calls; a periodic pass merges near-duplicate modes that share a single fix.
  3. The metric’s Insights card shows its failure modes sorted by frequency, defaulting to the last 24 hours — adjust the date range to widen or narrow the window.
These insights surface the dominant patterns behind a metric’s failures, so you can fix the agent instead of re-reading every flagged call.
Insights only investigate metrics you’ve enabled. To find misbehavior no metric measures, run Deep Research — a project-wide audit of a window of calls.

Viewing and generating insights in the dashboard

Navigate to Observability → Insights in the sidebar. The page displays a card for each metric you’ve enabled insights for. Use the View dropdown at the top of the page to filter by view. When a view is selected, the page shows only metrics and insights for agents in that view. Use the agents dropdown next to the date-range picker to filter by specific agents. When agents are selected, failure-mode counts, example calls, and failure rates reflect only calls handled by those agents; failure modes with no matching calls are hidden. Clear the selection to include every agent. Each card shows one of the following:
  • The latest failure-mode audit with identified themes
  • A message indicating not enough failures were found in the analyzed window
  • A Generate button if no audit has run yet
Click Generate (or Regenerate for an existing audit) to run a new analysis on demand. The page polls automatically until the audit completes. Each failure theme includes a title, a brief description, and clickable call IDs that link directly to the corresponding call logs.

Seeing the impact of an agent version

Every production call records the agent version that was live when it arrived, so each failure mode can be read against the agent changes around it. Expand a failure mode to see By agent version: the versions its calls ran on, newest version first, each with the date that version went live, how many of the mode’s calls it accounts for, and its own example calls. A mode whose calls sit almost entirely on the newest version is a regression that version introduced; one that thins out as versions advance is being fixed. Only versions that actually have calls in that failure mode are listed — a version with no failures there is absent rather than shown at zero. Calls that arrived before version capture existed are grouped into a single Version not recorded row, so the per-version counts always add up to the mode’s total. To compare a metric’s overall failure rate across versions, filter by version through the API: agent_version_ids on the failure-categories, category-calls, and failure-rates endpoints takes the version ids reported in each mode’s breakdown.
A version is recorded at the moment a call arrives, and never rewritten afterwards — editing your agent today does not move yesterday’s failures onto the new version.

Trend: did a release make it worse?

Click Trend on a failure mode to chart it over time. The line is the mode’s rate — its failing calls as a percentage of the calls evaluated for that metric — and the coloured band beneath it shows which agent version was live, so a step up that starts where a band changes is a regression that release introduced. Every point is one interval’s own rate, never a moving average. That matters: a window wide enough to smooth the line would blend across the release boundaries the chart exists to distinguish, and report rates no version ever ran at. Where traffic is thin, Cekura widens the interval each point covers instead — still a real measurement of a real period. Below the chart, each version gets a row: how many of the mode’s failures it produced, how many calls it had evaluated, its rate, and the percentage-point move against the version before it. That last column is the answer to “which version did this”. Two things the view deliberately will not do, because either would mislead:
  • A period with no evaluated calls is a gap in the line, not 0%. Before a metric was switched on for Insights — or through a window with no traffic — there is no rate to report, and drawing a 0% floor there would make the first real measurement look like a sudden regression.
  • Only versions that actually ran calls in the range are listed. A version that ran calls and produced none of this failure mode shows 0% — that is what fixing it looks like. A version with no traffic in the range is left out rather than shown at zero.
Agent versions belong to a single agent, so the trend compares one agent’s releases at a time. When a failure mode spans several agents, it opens on the one carrying most of its failures and offers the rest in a picker. The rate is measured over that agent alone — a busier sibling’s traffic in the denominator would let a mode look fixed purely because the traffic mix changed.

Enabling insights for a metric

Insights are opt-in per metric. A metric produces no failure modes until you turn it on and tell Cekura what a passing call looks like for it. On the Insights page, click Enable insights for a metric, pick the metric, then set its passing threshold — the condition a call must meet to count as a pass. Calls that don’t meet it are the failures Cekura investigates. The threshold uses the same condition shape as a rubric rule. It applies only to new evaluated calls after Insights is enabled: Cekura does not reprocess historical calls. Failed calls are triaged as they arrive, so the metric’s failure modes populate over time. To change the threshold later, reopen the metric’s card configuration; to stop analysing a metric, turn insights off.
Insights cost 10 credits per metric per day. The dialog shows this cost before you enable a metric.

Prioritizing metrics

Metric cards are grouped into Custom Metrics and Predefined Metrics sections. Drag the grip handle in a card’s top-right corner to reorder cards within their section, so the metrics your team watches most sit at the top. The order is saved for the whole project, so everyone viewing Insights sees the same arrangement. Reordering isn’t available to members with read-only Insights access.

Which metrics are eligible

Any metric you evaluate on call logs can drive insights — LLM Judge, code-based, and predefined metrics alike. Failure is decided by the metric’s own passing criteria rather than by its type, so metrics such as Latency, WPM, and Talk Ratio are no longer excluded. A metric needs both of the following:
  • Evaluated on call logs. A metric that doesn’t run in Observability can never produce a failure to investigate.
  • Insights enabled, with passing criteria set. Both come from the step above.
The project rubric no longer governs insights. Adding a metric to the rubric does not enable insights for it, and enabling insights does not make the metric count toward call success — the two are configured independently.