What Otis computes
Candidate causes
How Otis tests whether a deploy, a flag variant or a model change made things worse, and how the result reaches an insight.
A candidate cause is a tested link between one recorded change and a measure that got worse after it. Otis tests each new deploy, feature flag variant and model change for a regression. When a test finds one, Otis offers it as a possible explanation while it investigates related problems.
The tests ask whether a change made a measure worse. They don't report improvements. Otis can still mention a recent change next to an improvement, as described below, but no test stands behind that link.
This page explains what Otis tests and what the result does and doesn't show. Otis never presents a candidate cause as a proven cause, and the wording in an insight reflects that.
Purpose of change tests
A list of what changed doesn't say which change mattered. A regression that appears on the day of a deploy may have nothing to do with it. Otis compares the users who got the change with users who didn't, so that an insight can say whether the timing is backed by a measured difference.
Tested changes
Otis tests three kinds of change:
- Deploys. A new version of a service in an environment.
- Feature flag variants. A variant of a flag that users have newly been assigned.
- Model changes. A service that starts calling a different AI model.
Otis records two other kinds of change and doesn't test them. An edit to a measurement's definition and a break in telemetry are both changes in how things are measured. Otis treats them as reasons that two periods can't be compared, and never as explanations for a shift in behavior.
Comparison groups and measures
For each change, Otis forms two groups of users and compares them.
| Change | The two groups |
|---|---|
| Deploy | Users on the new version and users on the version before it. |
| Flag variant | Users assigned the new variant and users assigned the control. The control is a variant named control, or the largest other variant. |
| Model change | Users whose calls went to the new model and users whose calls went to the old one. |
Where the two versions or models ran at the same time, Otis compares the groups over that shared period. Otherwise it compares the period after the change with a period of the same length before it. The second kind of comparison is weaker, because anything else that changed between the two periods affects it too.
Otis compares four measures: frustration, errors, task failure and latency. Each is worked out for each user, so that one very active user doesn't dominate a group.
Result thresholds
Otis reports a regression on a measure only when each group has enough users, the difference is large enough to matter, and the difference is unlikely to be coincidence. It adjusts for having tested several measures at once.
Each result has a confidence of high, medium or low. A comparison over the same period earns more confidence than a before-and-after comparison. Large groups of similar size earn more than small or lopsided ones.
Limits of a candidate cause
- It isn't proof. Otis compares two groups from ordinary use. It doesn't assign users to the groups at random, unless your flag system does, so the groups can differ in other ways.
- No result doesn't mean no effect. Each test asks whether the new group is worse on a measure. A change that improved a measure produces no result, and neither does one with too few users to compare.
- Changes that land together aren't separated. Each change is tested on its own. If a deploy and a flag variant arrive on the same day, Otis doesn't work out which one is responsible.
- A flag at 100% has no comparison. With every user on the new variant, there is no control group, and Otis skips the test.
Candidate causes in an insight
A candidate cause never appears by itself. When Otis investigates a finding about errors or frustration, it is given the recent candidate causes, with the comparison and its confidence. It then decides whether each one fits the finding.
Otis is instructed to describe a cause it accepts as a likely candidate, or as consistent with the finding, and not to state that the change caused the problem. It is also instructed to say so when it considered a recent change and judged it unrelated.
Changes mentioned without a test
Otis also sees the list of recent changes whenever it investigates, whatever the finding is about. It can set any movement against a change on that list, including an improvement, in words such as "the rise followed Monday's deploy". That sentence reports timing. It carries no comparison of two groups, so give it less weight than a candidate cause that states the groups and the difference between them.
Candidate causes in Otis
You read candidate causes in the text of an insight. You can also ask Otis which changes were recorded around the time of a finding. Changes describes that.
Reading a claim about a change
- Look for the comparison. A sound claim names the two groups, their sizes and the difference between them.
- Prefer a comparison over the same period. Two versions measured at the same time control for the day of the week and for other events. A before-and-after comparison doesn't.
- Use a flag to settle it. If your flag system assigns variants at random, the two groups differ only by the variant. That is the strongest evidence Otis can give that a change caused a difference.
Related
- Changes covers how Otis records deploys, flag variants and model changes.
- Feature flags covers how your app reports flag variants.
- Insights covers how a finding is investigated and written.
- Signals covers the frustration and error signals the tests read.