What Otis records
Tasks
What a task is, how Otis finds tasks inside a session, and how it decides whether each one worked.
A task is one piece of work that a user set out to do, such as drafting an email, fixing a failing test, or exporting a report. Otis divides each session into tasks. For each task, it records what the user was trying to do and how the attempt turned out.
Otis also groups tasks that share a goal into intent clusters. This page explains where the boundaries of a task come from and what its outcome means, so that you can read task counts and outcomes correctly.
Purpose of tasks
A session covers one whole visit, and a visit usually includes several goals. A user might ask for a summary, rewrite a paragraph, and then export the result. One of those goals can go well while another goes badly, and a single outcome for the whole session hides the difference.
A single message or a single action is too small a unit, because on its own it rarely shows whether the user got what they wanted. A task is the smallest unit that has both a goal and a result, so it is the unit Otis scores.
Task detection
Otis uses three methods, depending on what a session contains. Each method produces its own kind of task.
Conversations with your AI
When a session contains messages between a user and your AI, Otis reads the user's messages in order. At each message, it decides whether the user is continuing the same goal or starting a new one. A new goal starts a new task.
Otis starts a new task only when the change of goal is clear. A message that could be read either way stays in the current task, because splitting too readily would break one piece of work into fragments.
Two rules affect where the boundaries fall:
- Separate conversations stay separate. If your app sends a chat ID, Otis divides each chat into tasks on its own. Two chats that are open at the same time never merge into one task.
- Only the user's own messages can start a task. An agent often sends prompts of its own to a model while it works. Otis keeps those calls inside the task they belong to, and they never start a new one.
Activity in your product
A session can also contain actions that involve no AI message, such as opening a record, saving a draft, or sharing a link. Your app records these with custom events and traced functions.
Otis groups these actions into bursts of activity. A burst ends when the user pauses for a few minutes, and each burst with more than a couple of actions becomes one task. Otis names the task after the action that occurs most often in it, such as prospect.view. If no single action occurs most often, the task is named mixed_activity.
Some events fire on a timer and don't show that the user did anything, such as a status check that runs every 20 seconds. Otis recognizes these background events, and leaves them out when it measures idle time and when it counts actions.
Tool calls from an agent
You can instrument a Model Context Protocol (MCP) server or a command-line tool that AI agents use. In that case Otis sees the calls an agent makes to your tool. It doesn't see the conversation between the agent and the person directing it.
Each call may carry a short statement of what the agent is trying to do. Otis groups consecutive calls that share the same statement into one task, and a call with no statement stays in the task around it. A long pause between calls also starts a new task. When the agent gives no statement at all, those pauses are the only boundaries.
Sessions that mix conversation and activity
Many sessions contain both messages and product actions. Otis then produces both kinds of task for the same session. Conversation tasks are built from the messages, and activity tasks are built from the actions. The two kinds never overlap, so a message is never counted in an activity task and an action is never counted in a conversation task.
An action that a user takes in the middle of a conversation, in the same part of your product, stays with the session and isn't counted in a task.
Separation by surface
A surface is a named part of your product that users work in, such as one copilot, one MCP server, or one command-line tool. If your app names its surfaces, Otis finds tasks within each surface separately. A conversation with a chat ID is the exception. Otis keeps it together even if its messages carry different surface names.
Name the surface on all of the telemetry for a flow, or on none of it. If only part of a flow is labeled, Otis treats the labeled and unlabeled parts as different surfaces and splits the activity between them.
Task timing
Otis divides a session into tasks once the session has been idle for 30 minutes. The wait gives late telemetry time to arrive, and it keeps Otis from closing a task while the user is still working. If the user returns to the same session later, Otis divides the new activity in the same way.
Otis only divides sessions that are large enough to analyze. It skips a very short session, and one with only a few spans. A span is one recorded operation, such as an AI call, a traced function, or a custom event.
The task record
- Intent. A short summary of what the user was trying to do, and a category that names the kind of work. For a conversation task, Otis writes the summary from the messages. For an activity task, the summary lists the actions in order. For a tool task, the summary is the statement the agent gave, or a list of the tools it called if it gave none. Intent covers both.
- Outcome. One of four values, described in the next section.
- Signals. What the user's messages expressed, such as frustration, confusion or delight, along with any feedback your app sent. Signals lists them.
- Size. The number of user messages, the number of spans, and how long the task lasted.
- Errors. How many spans failed, and the causes.
- Cost. The token usage and cost of the AI calls in the task.
- Context. The user, session, chat, document and surfaces that the task belongs to.
Task outcomes
Every task has one of four outcomes:
| Outcome | Meaning |
|---|---|
| Success | Otis found evidence that the user got what they wanted. |
| Struggled | Otis found evidence of a problem, and no evidence of success. |
| Abandoned | The user stopped before the work was finished. |
| Unknown | Otis found no evidence either way. |
Otis looks for evidence in the following order and uses the first kind it finds.
- An outcome your product reports. This applies to activity tasks. An event that your funnel definition lists as a success stage, such as
share, marks the task as a success. A failure stage, such asdiscard, marks it as struggled. If one task contains both, the failure counts. Funnels and artifacts describes how to send these stages. - What the user said or rated. In a conversation task, Otis reads each user message for signs of how the work is going. Delight, or a statement that the AI helped, marks the task as a success. So does an action that Otis recognizes as the user putting the result to use, such as saving, exporting or copying it. Frustration, confusion, or a request to reach a human marks it as struggled. Feedback that your app sends, such as a thumbs-up or a correction, counts in the same way. So does the feedback an agent files about a tool. If a task shows both satisfaction and frustration, Otis records it as a success.
- How the work ended. If the last operation in the task failed, the task is struggled. If the user cancelled or interrupted it, the task is abandoned. If it returned a successful HTTP status, the task is a success. Otherwise, an earlier failure that the task never recovered from marks the task as struggled.
- A very short conversation. A conversation task that ends almost at once and shows no other evidence is abandoned.
If none of these applies, the outcome is unknown.
Reading outcomes
- Unknown means no evidence. Otis doesn't guess an outcome. A task with no feedback, no outcome event and no error stays unknown, and some share of unknown tasks is normal. Sending feedback and artifact stages gives Otis more evidence and reduces that share.
- A struggled task may still have been finished. Struggled records that Otis saw a problem and saw no sign of success afterward. The user may have completed the work without saying so.
- Compare tasks of the same kind. A conversation task usually gets its outcome from what the user wrote. An activity task usually gets it from an event or an error. A success rate that combines the two kinds mostly reflects how many of each there are, so compare rates within one kind.
Tasks in Otis
The data browser has a list of tasks that you can filter by type and outcome. Opening a task shows the spans it contains. Each session page also lists that session's tasks, and a number in an insight can link to the tasks behind it.
Related
- Sessions covers the unit that Otis divides into tasks, and Sessions in the SDK guide covers session and chat IDs.
- Signals and Feedback cover the evidence Otis reads from messages and from explicit reactions.
- Errors covers what Otis records when an operation fails, and which failures count against a task.
- Funnels and artifacts covers the stages your app can send to report how a piece of work ended.