How Otis knows your product
Statistical tests
The statistical rules Otis applies before it reports a difference, what they protect against, and what they don't.
Before Otis reports that one group of users differs from another, or that a measure has moved, it tests whether the difference could plausibly be chance. This page describes the rules it applies.
It also lists what those rules don't cover. Otis works from ordinary usage data, so every finding describes an association. The rules reduce the chance of a false alarm. They can't show that one thing caused another.
Purpose of the tests
A product with many users, many events and many ways to group them will always contain differences that look striking and mean nothing. A system that reported each one would bury the real findings. The rules below are how Otis decides that a difference is worth an investigation.
Sources of candidates
Otis's detectors look for four kinds of pattern, and each is tested before it becomes a candidate:
- A group that differs. One group of users compared with everyone else. Examples are the users with a given property, or the accounts on one plan.
- A problem that concentrates. One operation, task category, plan or funnel step that carries a higher rate of a bad outcome than all the others together.
- A population that moved. A change in the number of active, paying, quiet or converted accounts between two dates.
- A rate that shifted. A change in an error rate, a funnel's success rate or a cancellation rate between two adjacent periods.
Some candidates are chosen without a statistical test: a recurring problem shared by several users, a session that stands out, and an open-ended exploration. An insight from one of these rests on the investigation that follows, and on the counts it reports.
Rules for a tested candidate
Enough data
Each side of a comparison needs a minimum number of users or accounts, and a minimum number of cases of each outcome. A comparison in which almost nobody, or almost everybody, has the outcome can't be read.
A project with few accounts therefore gets no comparison between groups of accounts.
Strong evidence
Otis computes a range of plausible values for the difference, and requires that the whole range lies on one side of zero. It sets the range at a stricter level than the 95% used in most analytics tools.
A difference large enough to matter
Otis separates a difference that is real from one that is large. It counts a difference as meaningful only when it is large in absolute terms, large relative to the baseline rate, and affects enough cases. A difference that is clearly real but smaller is reported as a direction only. A difference that is shown to be small is dropped.
Many tests at once
A detector tests many groups, and the more it tests, the more likely one passes by chance. Otis corrects for this within each day's tests, including the groups it considered and didn't report. It corrects again as tests accumulate across days.
A candidate that fails a correction never reaches an investigation.
People, not events
One very active user can produce hundreds of tasks. If each task counted as an independent case, that user's habits would look like a pattern.
For most comparisons, Otis treats the user as the unit, and treats the account as the unit when the group is defined by an account property. Many tasks from one person then count as one source of evidence.
Like with like
For some comparisons, Otis compares each group with otherwise similar users, such as users on the same plan and of similar age in your product. If there are too few similar users, Otis loosens the match and says so in the candidate.
Timing
- No test in the first week. Otis runs no statistical test on a project until it has seven days of data. After that it tests at set intervals, and not every day.
- Unfinished work isn't counted as failure. Funnel rates are read once each group of entries has had time to progress, so recent entries that are still in progress don't read as a drop.
- Weekends are balanced. A comparison of two adjacent periods is refused when one has noticeably more weekend days than the other.
- Otis doesn't compare across a break. Several comparisons are refused when the rate at which your telemetry is sampled changed inside the period, or when the data is stale.
Comparisons that only restate a definition
Otis excludes groupings that would make a finding circular. A group that is defined by the outcome it is compared on, for example, says nothing about that outcome.
Rules for the investigation and the text
A tested candidate is then investigated and written up, and further rules apply. Checks before publishing lists them. Three bear on how to read a claim:
- A comparison must state its unit. Otis must say whether a rate counts users who ever did something, events, or an average for each user.
- Causes are not asserted. Otis describes an association unless it has a comparison that supports more, and it names the most obvious alternative explanation.
- A mechanism is marked as observed or inferred. When the explanation for a pattern comes from what users wrote and not from a measurement, the insight may say only what users report.
Findings that turn out wrong
Otis doesn't retract a published insight. Three things can happen instead:
- A later measurement replaces it. The new figures are published as a new version, and the earlier version stays readable.
- It goes stale. An insight that isn't found again within 30 days moves under "Resolved".
- You mark it as wrong. Otis then treats the claim as rejected when it next meets it. Acting on an insight describes this.
Limits of the rules
- An unmeasured cause. Your users weren't assigned to the groups at random, so anything that differs between two groups can explain a difference.
- A wrong question. The rules test whether a difference is real. They don't test whether it is the difference that matters.
- Indirect measures. Frustration, confusion and similar outcomes are judgments that an AI model makes about a message. They have their own error rate, which the tests don't include.
- Small projects. A project below the minimums gets few tested findings, however large its differences look.
Test results in Otis
Otis doesn't show the test results directly. An insight shows:
- A confidence level of high, medium or low. It summarizes how much evidence the investigation gathered, and it isn't a statistical result. Reading confidence explains it.
- The number of cases the finding was measured on.
- "Still open", which states the limits of the finding in its own terms.
If your team turns on Show the numbers in its guidelines, Otis adds ranges of uncertainty to the figures it reports.
Related
- Insights covers the investigation and the checks on what Otis writes.
- Candidate causes covers the separate test Otis runs on each deploy, flag variant and model change.
- Evidence covers how quotes and recurring problems are chosen and checked.
- Segments and cohorts covers the groups that comparisons are made between.
Context
Which of the things Otis knows about your product reach each part of its work, what they can change, and what they can't.
Data health and gaps
What Otis accepts and drops when telemetry arrives, how it notices a break, what it waits for in a project's first days, and how long it keeps each kind of record.