OtisDocs

What Otis records

Errors

What Otis records when an operation fails, how it names the cause, and which failures count against a task.

An error is an operation that failed. For each failed span, Otis records what the failure was and gives it one cause from a fixed list, such as a rate limit, a timeout, a tool that returned an error, or a user who cancelled.

Otis uses the cause for two decisions. The first is whether the failure counts toward your error rates. The second is what the failure means for the task it happened in. This page explains both, because they don't always agree.

Purpose of error causes

A failed status says that something went wrong. It doesn't say whether your product is at fault. A provider outage and a user who pressed stop both end an AI call early, and they call for different responses. The cause keeps them apart, so that an error rate measures failures of the product and leaves out the choices users make.

The error record

  • The type, message and stack trace of the error, when the SDK has them. The message and the stack trace are scanned for personal data in the SDK and again when Otis receives them, so they are stored in redacted form.
  • One cause. The next section describes how it is chosen.

A function that catches an error and returns normally is recorded as a success. Make failures visible describes how to report a failure that your code handles.

Cause assignment

The SDK names the cause when it knows it. It classifies a failed AI call from the error's name and status code, and it knows when an MCP tool returned an error or a command-line run was interrupted.

When the SDK names no cause, Otis works one out from the span. It checks the following in order and uses the first that applies: an MCP error code, an interrupted command-line run, a command that exited with a non-zero code, a recorded exception, an HTTP 5xx status, an HTTP 4xx status, and a failed status with nothing else to go on.

The list has 19 causes in seven groups: AI calls, tools, command-line runs, HTTP, exceptions in your own code, cancellation by the user, and a failed status with no other detail. Error causes lists each one.

Two causes are recorded on spans that didn't fail in the usual sense:

  • A refusal. When a model declines to answer, for example because a content filter stopped it, the span keeps a success status and gets the cause ai_content_filter. A refusal counts toward error rates.
  • A cancellation. When the user stops a response while it is streaming, the span gets the cause cancelled.

Errors that count

Otis applies two separate rules to a cause.

CauseCounts toward error ratesEffect on the task
Most causesYesThe task failed
http_4xxNoThe task failed
cancelledNoThe task was abandoned
interruptedYesThe task was abandoned

An HTTP 4xx status is usually the caller's mistake, so it stays out of error rates. A cancellation is the user's choice, so it stays out too. An interrupted command-line run does count, because a user usually interrupts a run that is stuck.

Error totals

  • On a task. Otis counts the spans in the task with a cause that counts toward error rates, and records how many there were of each cause. The cause of the task's last step, and any failure earlier in the task, also help decide its outcome. Tasks describes the order.
  • On a session. Otis keeps the same counts over all the spans in the session.
  • On an operation. For each day, Otis counts the failures of each operation and the users who hit them, by cause and by surface. It uses these counts to find the operations where failures are concentrated, and the ones where the share of affected users is rising or falling.

Recurring errors

Otis groups failures that look like the same underlying error. It identifies one in the first of these ways that the span supports:

  1. The error type and the top stack frames from your own code, with line numbers ignored.
  2. The error type and the message, with the variable parts removed, such as IDs and long quoted strings.
  3. The error type and the cause.

The first way needs a JavaScript stack trace. Errors from other runtimes are grouped by message.

Otis rebuilds these groups daily from recent failures, and keeps a group that several different users hit. It can raise the most widespread groups as insights.

Errors and the error signal

The signal error_or_tool_issue is a different record. It fires when your AI's reply says that something failed. An error is a failed operation in your telemetry. Either can occur without the other, and the signal never changes a task's outcome.

Reading errors

  • A handled error is invisible. If your code catches a failure and carries on, Otis sees a success unless you report the failure.
  • Error rates and outcomes use different rules. A task can be struggled because of an HTTP 4xx status that doesn't appear in any error rate.
  • A cancellation is not a failure. A task that ends with a cancelled response is abandoned, and the cancellation isn't counted as an error.
  • Evidence of success comes first. If the user expressed satisfaction during a task, the task is a success even when its last step failed.

Errors in Otis

In the data browser, the span list has an Error column and a filter by cause. A cause is shown in red only when it counts toward error rates. The task list and the session list each have an Errors count, and opening a task or a session shows the count for each cause.

  • Errors and exceptions covers what the SDK records automatically, the full list of causes, and how to report errors from your own code.
  • Tasks covers how a failure affects a task's outcome.
  • Signals covers error_or_tool_issue, which records what the AI said.

On this page