OtisDocs

SDK Reference

Privacy

PII redaction and identifier hashing

Otis aims to never hold Personally Identifiable Information (PII), since it has little to no value in product analytics.

Otis applies two independent privacy mechanisms, each with a client-side and a server-side layer:

  • PII redaction — scrubs PII content from span text before it reaches Otis.
    • Client-side runs in the SDK before spans are exported. Regex patterns catch structured PII (credit cards, emails, phone numbers, SSNs, API keys, JWTs, etc.).
    • Server-side runs when Otis receives spans and before they touch persistent storage. A machine-learning model catches unstructured PII that regex can't detect (names, addresses, organizations, other contextual PII) and composes with the client-side layer.
  • Identifier hashing — pseudonymizes raw user, session, and group IDs so raw values never reach analytics storage.
    • Client-side HMAC applied by the SDK before values are exported, using a seed taken from your API key. Raw identifiers never travel over the wire. Set identifierHashKey if this layer has to stay opaque to Otis.
    • Server-side HMAC applied when Otis receives spans and before they touch persistent storage, using a separate per-project secret held by Otis. Double-hashing defeats rainbow-table attacks on the space of likely raw identifiers.

Each layer is independent defense-in-depth against the previous one failing or missing something.

Client-side PII redaction

Enabled by default. 40+ regex patterns with Luhn and similar validators to minimize false positives. Entity linking ensures the same PII value gets the same placeholder within a span.

initOtis({
  serviceName: "my-app",
  piiRedaction: {
    enabled: true,                    // default
    disabledPatterns: ["ipv4"],       // skip specific patterns
    customPatterns: [{                // add custom patterns
      name: "internal_id",
      regex: /INT-\d{10}/g,
      placeholder: "[INTERNAL_ID_N]",
      priority: 50,
    }],
  },
});

Entity linking

Within a single span, the same PII value receives the same placeholder. Maps reset between spans to maintain context separation.

Input:  "Contact john@acme.com for help. CC john@acme.com for the team."
Output: "Contact {REDACTED_EMAIL_1} for help. CC {REDACTED_EMAIL_1} for the team."

Which attributes are scanned

AI prompt/response attributes, user-corrected text, and user/group property values are scanned.

Scanned attributes: ai.prompt, ai.prompt.messages, ai.prompt.lastUserMessage, ai.response, ai.response.text, gen_ai.input.messages, gen_ai.prompt, gen_ai.prompt.messages, gen_ai.output.messages, gen_ai.response, gen_ai.response.text, gen_ai.completion, ai.tool.input, ai.tool.output, ai.response.tool_calls, user_message, response_message

Scanned feedback fields: The comment and expected attributes in sendFeedbackSignal are scanned.

Scanned exception fields: exception.message and exception.stacktrace from sendException. exception.type (and its error.type twin) is not scanned — it is a low-cardinality error class name, so keep personal data out of your error type names. Home-directory paths in stack traces are redacted to the username segment only, so frames stay readable (/Users/jsmith/app/x.ts → /Users/{REDACTED_USERNAME_1}/app/x.ts).

Scanned prefixes: traits.*, metadata.*, otis.metadata.*, properties.*, session_properties.*

You never write the metadata prefix yourself — the SDK adds it to whatever you pass as metadata: metadata: { jobId } becomes otis.metadata.jobId on withContext, wrap, traced, sendException, and sendFeedbackSignal spans.

Inverse contract. Anything else — top-level attributes on sendEvent calls, OTel-semantic identifiers (user.id, session.id), structural fields (format, doc.id, tool.name), the artifact.id / artifact.stage / artifact.type set — is treated as identifier-shaped and ships unscanned. To route a freeform value through redaction without changing the namespace, use the opt-in segment rule below.

MCP servers: captured tool input/output is outside this scanned set, so the MCP instrumentation applies this client-side redaction to it directly (on by default) and offers captureInput / captureOutput controls.

Opt-in scanning with otis_redact_

Any attribute key with an otis_redact_ or sensitive_ segment prefix is value-scanned for PII, regardless of namespace. The match is case-insensitive and applies at any depth, and the . form works too (otis_redact.note).

otis_redact_ and sensitive_ are exact aliases. Prefer otis_redact_ in new code — it names the instruction rather than classifying the data. sensitive_ keeps working indefinitely.

sendEvent("document.export", {
  format: "pdf",                                   // not scanned
  sensitive_note: "Drafted with alice@acme.com",   // scanned → stored as `note`
});

sendArtifactEvent("doc-abc123", {
  stage: "share",
  sensitive_note: "Sent to alice@acme.com",        // scanned → stored as `artifact.note`
});
KeyScanned
sensitive_noteyes
sensitive.noteyes
artifact.sensitive_noteyes
hello.world.sensitive_emailyes
Foo.SENSITIVE.baryes (case-insensitive)
otis-redact-noteyes (hyphen is a separator)
otis_redact.noteyes (trailing . is fine)
nonsensitive_thingno (no segment boundary)
my-sensitive-noteno (no segment boundary — a hyphen doesn't create one)
email.sensitiveno (prefix the parent, not the leaf)
otis.redact.noteno — inside otis_redact the separator must be _ or -
sensitiveNoteno — camelCase is not a marker

The same rule applies on the server-side layer, so opting in covers both regex and NER passes.

The marker needs a separator, and inside otis_redact it must be _ or -. Write otis_redact_note, otis-redact-note, otis_redact.note or sensitive.note. otis.redact.note is not a marker — a . there reads as a path segment, not part of the token, so the field gets no protection. sensitiveNote and otisRedactNote are likewise ordinary field names, not directives.

This is deliberate rather than an oversight, and it is the one place Otis does not accept the camelCase spelling. Elsewhere it does — firstName is dropped by the same rule as first_name — because there both spellings name the same field. A marker is different: it renames the attribute. If sensitive + a capital letter were a marker, then sensitiveContent, sensitiveData and sensitiveFields would become directives, and your sensitiveContent field would arrive in Otis called Content. Those are ordinary names, so the marker requires a separator to stay distinguishable from one.

The sensitive_ prefix is removed from the stored attribute. It's a flag that tells Otis to scan the value, not part of the attribute's name — so once the value has been scanned, the prefix is stripped and the attribute is stored under its plain name (sensitive_note → note, artifact.sensitive_note → artifact.note). Flag the value; query the clean name. (If the plain name already exists on the same event, the flagged key is left as-is to avoid overwriting it.)

A flagged key that names a PII field is dropped, not just scanned. Otis drops property keys whose name reveals PII (email, ssn, first_name, …) even when the value is redacted, because the key name alone is revealing. The opt-in marker is optional here too: otis_redact_email is dropped exactly like email, and sensitive_email behaves identically — the two spellings are exact aliases everywhere. Use it to protect a value under a non-PII-named key (like otis_redact_note); it won't rescue a key that already names a PII field. To keep such a key, use the opt-out below.

Opting out with otis_dont_redact_

Any attribute key with an otis_dont_redact_ (or .otis_dont_redact.) segment prefix is exempt from everything — key-based dropping, the client-side regex, the server-side regex, and the server-side model. The value is stored exactly as you sent it, and the marker is stripped from the key before storage.

setUserProperties({
  email: "ada@acme.com",                    // dropped — the key names a PII field
  otis_dont_redact_email: "ada@acme.com",   // kept verbatim → stored as `email`
  otis_dont_redact_zip: "94114",            // kept verbatim → stored as `zip`
});

Same matching rules as the opt-in: case-insensitive, any depth, boundary is ^ or . — so my_otis_dont_redact_note does not match, and neither does the camelCase otisDontRedactNote. Inside otis_dont_redact the separator must be _ or - (otis-dont-redact-email works, otis.dont.redact.email does not); a trailing . is fine (otis_dont_redact.email).

otis.dont.redact.email is not the opt-out, and for a PII-named field it does the opposite. With dots inside the token it is read as an ordinary key called email, so it is dropped by name — the outcome the marker exists to prevent. Use _ or - inside the token.

Two behaviour changes on upgrade.

The opt-out now honours hyphens. otis-dont-redact-<field> did not previously register, and for a field whose name Otis drops by default it produced the opposite of what it asked for: properties.otis-dont-redact-email was dropped entirely. It is now honoured, so a value you were losing will start being stored verbatim — which is what the marker asks for, but check that the fields you spell this way are ones you intend to keep in the clear.

The opt-in now honours hyphens too, which can capture an existing field. A key like properties.sensitive-data or properties.sensitive-content was previously an ordinary attribute; it now reads as the marker plus a field name, so its value is scanned and it is stored under the shortened name — properties.data, properties.content. Queries and dashboards on properties.sensitive-* will stop matching, with no error. (sensitive_data has always behaved this way; the change extends it to the hyphenated spelling.) If you have a field genuinely named sensitive-<something>, rename it or use otis_dont_redact_ on it.

This transfers responsibility for the field to you. The value is written to storage exactly as sent, with no key-drop, no regex and no model. Never put it on a field that can carry a natural person's data.

When you actually need it. Redaction is irreversible and happens before storage, and both removal paths have false positives you cannot otherwise escape. Key-dropping matches a fixed list of well-known names, so a legitimate address field meaning a wallet address — or a zip meaning a compression format — is dropped by name. Value scanning is a model, and models are wrong sometimes.

Use the narrowest thing that works: prefer renaming the field (wallet_address rather than address), then the opt-out on a single attribute. Avoid applying it across a whole namespace.

KeyResult
otis_dont_redact_zipkept verbatim → stored as zip
otis_dont_redact_emailkept verbatim → stored as email (key-drop bypassed)
otis_dont_redact.notekept verbatim → stored as note
OTIS_DONT_REDACT_notekept (case-insensitive)
my_otis_dont_redact_notenot matched — normal rules apply

Property key dropping

Property keys set via setUserProperties and identifyUser are checked against a curated list of well-known PII field names (e.g. ssn, email, first_name, phone_number, date_of_birth, credit_card, address). Matching entries are dropped entirely, since the key name itself reveals PII intent even if the value is redacted.

Hyphens are normalized to underscores, so first-name matches first_name. camelCase is matched too, so firstName, emailAddress, phoneNumber, dateOfBirth, creditCard, accountNumber and ipAddress are all dropped by the same entries as their snake_case spellings. Suffix matching catches compound keys like customer_ssn or primary_email.

camelCase matching arrived in a later SDK version. If you were relying on a property like firstName or accountNumber reaching Otis — for example as an opaque join key rather than as personal data — it will start being dropped on upgrade. Rename it, or use the otis_dont_redact_ opt-out to keep it.

camelCase is matched as a whole key only: firstName is dropped, but customer_firstName is not (its snake_case spelling customer_first_name is, via suffix matching). Spell compound keys consistently — all snake_case or all camelCase — rather than mixing the two in one name.

Group properties are governed server-side

setGroupProperties describes an account or company — not the individual person that PII redaction protects — so whether group properties are redacted is a server-side, per-account-type policy, not a client-side rule. The SDK does not key-drop or value-scan setGroupProperties calls; it defers them to Otis. By default the server redacts group properties just like user properties, but you can keep firmographics (e.g. account name, plan, industry) for specific group types from Settings → Ingestion → PII Redaction. The otis_redact_ opt-in still applies client-side to any group property value you explicitly flag.

initOtis({
  serviceName: "my-app",
  piiRedaction: {
    dropPIIPropertyKeys: true,                   // default: true
    additionalPIIPropertyKeys: ["employee_id"],  // merge with defaults
  },
});

Property values not dropped by key matching are still regex-scanned. For example, properties.notes = "Call 415-555-1234" redacts the phone number.

Use debug: "filter" to see which keys are dropped:

[otis:filter] { piiKeyDropped: 'properties.email' }

Override scanned attributes

scanAttributes and scanAttributePrefixes replace the defaults when provided:

initOtis({
  serviceName: "my-app",
  piiRedaction: {
    scanAttributes: ["custom.user_input", "custom.response"],
    scanAttributePrefixes: ["user_data."],
  },
});

To extend defaults rather than replace, spread the exported constants:

import {
  initOtis,
  DEFAULT_SCAN_ATTRIBUTES,
  DEFAULT_SCAN_ATTRIBUTE_PREFIXES,
} from "@runotis/sdk";

initOtis({
  serviceName: "my-app",
  piiRedaction: {
    scanAttributes: [...DEFAULT_SCAN_ATTRIBUTES, "custom.user_input"],
    scanAttributePrefixes: [...DEFAULT_SCAN_ATTRIBUTE_PREFIXES, "user_data."],
  },
});

Server-side PII redaction

Enabled by default on every project environment. A second pass runs on the ingest collector before data is persisted, using a machine-learning model to detect unstructured PII that regex patterns can't reliably catch: personal names, street addresses, organization names, and other contextually identified entities. The server-side pass runs on the same set of attributes as the client-side layer.

What it adds on top of client-side redaction

  • Natural-language entities. Names, addresses, and organizations that don't match any regex pattern.
  • Cross-attribute entity linking. The same person name appearing in user_message and ai.response.text within the same span receives the same placeholder.
  • Composition with client-side redaction. If the SDK has already redacted part of a string (e.g. replacing an email with {REDACTED_EMAIL_1}), the server-side layer still analyzes the surrounding text and can detect adjacent PII the SDK missed, without producing overlapping or conflicting replacements.
  • Always-on regex union. The same structured-PII regex patterns the SDK uses also run server-side, so a client with PII redaction disabled is still protected.

Configuration

Server-side redaction is configured per project environment from the Otis UI (Settings → Ingestion → PII Redaction). Enabled by default; the other settings below have sensible defaults and rarely need to change.

SettingDefaultEffect
EnabledOnRun server-side redaction on every span for this environment
Confidence thresholdTuned for balanced precision/recallMinimum model confidence for an entity to be redacted (higher = fewer false positives, possibly lower recall)
Disabled entity typesNoneEntity types to skip (e.g. keep organization names if they're never sensitive for your app)
On errorpassthroughpassthrough — keep the span unredacted if the redaction model is unavailable
drop — discard the span if redaction can't run
Redact group properties by defaultOnWhether setGroupProperties values are redacted for group types you haven't set explicitly. Turn off to retain account firmographics (name, plan, …) that aren't PII for a B2B account
Per-type overridesNoneSet a specific group type to redact or keep, overriding the default for that type — e.g. keep company while everything else stays redacted, or (with the default off) redact only household

Group properties

Group properties (setGroupProperties) describe an account, not a person, so the firmographics a B2B product attaches to an account — company name, plan, seat count, industry — are usually not PII with respect to that account. Group-property redaction is decided here, on the server, per group type: a default for types you haven't set, plus explicit per-type overrides that win over it.

  • Default on + keep company → company accounts keep their properties; every other group type is still redacted.
  • Default off + redact household → all group types pass through except household, which holds real people.

Each per-type setting says directly whether that type is redacted, so its meaning never depends on the default. A sensitive_* group property value is always scanned regardless of these settings, so you can flag the occasional field that does carry end-user PII (e.g. sensitive_account_owner) and keep it protected. With the defaults (redact on, no overrides), group properties are redacted exactly like user properties.

Turning server-side redaction off is supported but not recommended. Client-side redaction alone will not catch unstructured PII like names or addresses.

Behavior on failure

If the server-side model is unavailable, the regex union still runs, so structured PII (credit cards, SSNs, emails, phone numbers, etc.) is always caught. The on error setting determines what happens if both layers fail: passthrough accepts some risk of PII reaching storage in exchange for no data loss; drop discards the span.

Keep client-side on too

Even with server-side redaction enabled, keep client-side redaction on. Client-side catches structured PII before it leaves your infrastructure, which reduces your surface area in the event of a transport-layer issue or misconfiguration. The two layers are designed to compose.

Identifier hashing

All identifiers — user IDs, session IDs, and group IDs — are pseudonymized via HMAC-SHA256 before they leave your application. This lets you correlate activity for analytics without exposing raw identifiers (emails, internal user IDs, etc.) in analytics storage.

Enabled by default

Identifier hashing is enabled by default. The HMAC seed comes from your API key, so no additional configuration is needed.

initOtis({
  serviceName: "my-app",
  // Identifier hashing is enabled by default.
  // To disable (not recommended): identifierHashing: false
});

Hashing happens at entry points (contextFromChatRequest, withContext, wrap context, identifyUser), so downstream code always sees already-hashed values.

Identities stay the same across API keys

Each environment (production, development) has one identity seed, and every API key carries it:

sk-otis-K7mQ2xVb9TnR4wYc8LpZ3a-h5Gd1SfJ6uNe0BkXoW2tMq
└─────┘ └────────────────────┘ └────────────────────┘
 scheme          seed                random part

The first key you create for an environment fixes the seed, and every key created after it carries the same one. Rotating a key, or giving each service its own, therefore does not change who your users are: one person hashes to one ID under every key of the environment. Only the random part differs between keys, and it is what authenticates.

Keys in this format need @runotis/sdk 0.8.0 or later in every service that uses them, including the browser bundle. An older SDK does not know the format: it reports the same people under IDs of its own, which change when it is upgraded.

Replacing a key created before identity seeds

An older key, sk-otis-{project id}-{48 hex characters}, has no seed, so Otis cannot tell what identities it produces. It keeps working. When you create its replacement in Settings → Ingestion, the dialog asks for the key you are replacing. Paste it, and the new key continues its users, groups and sessions.

The dialog offers this once per environment. If you create a key without pasting one, later keys continue that new key instead, and the only way to keep the old identities is identifierHashKey below. A key revoked more than 30 days ago can no longer be pasted either.

Hashing with a secret of your own

identifierHashKey replaces the API key as the source of the seed:

initOtis({
  serviceName: "my-app",
  apiKey: process.env.OTIS_API_KEY,
  identifierHashKey: process.env.OTIS_IDENTIFIER_HASH_KEY,
});

It has two uses:

  • Keep the identities of an earlier key. Set it to that key, even one you have revoked, and identifiers hash as they did under it whatever key you authenticate with.
  • Keep the seed away from Otis. Set it to a long random secret that you generate and hold. Otis never receives it, so the client-side hash stays a boundary Otis cannot cross. See Server-side hashing layer.

Every instance that should agree on who a user is must carry the same value: your server, your browser bundle, each service. An instance without it hashes with its API key and reports the same people under different IDs. Server runtimes read OTIS_IDENTIFIER_HASH_KEY when the option is not passed; browsers have no environment, so pass it explicitly there.

A wrong value here is not caught by anything downstream: your API key still authenticates, spans still arrive, and every user is reported under a new ID. So initOtis throws if identifierHashKey starts like an Otis key but is not a complete one, which catches a truncated paste or a key copied with its quotes. It cannot catch a single wrong character. Where the SDK is disabled, or has no API key, it sends nothing, so the same fault is only a warning.

Upgrading from an SDK before 0.8.0

Earlier SDKs hashed client-side only when the key was passed as the apiKey option. An SDK that found its key in the OTIS_API_KEY environment variable sent raw identifiers and relied on the server-side hash alone. From 0.8.0 both hash the same way.

If your server is configured through OTIS_API_KEY with no apiKey option, upgrading changes the user, group, session and artifact IDs it reports, once. To keep the existing IDs instead, set identifierHashing: false, which is what that configuration was doing. Installations that pass apiKey are unaffected.

What gets hashed

InputOutput prefix
User IDusr_v1_<base64url-hmac-sha256>
Session IDses_v1_<base64url-hmac-sha256>
Group IDgrp_v1_<base64url-hmac-sha256>

HMAC input is domain-separated, so the same raw string passed as a user ID vs a session ID produces different hashes.

Helpers to check whether a value is already hashed

import { isHashedUserId, isHashedSessionId, isHashedGroupId } from "@runotis/sdk";

Already-hashed values pass through without re-hashing, so it's safe to call identifyUser(hashedId) in code paths where the ID might have already been hashed upstream.

A value counts as already-hashed when it carries the prefix and a 43-character base64url body — the shape the SDK itself produces. A prefixed value with any other body is treated as plaintext and hashed. Passing a hand-constructed usr_v1_… of your own devising therefore hashes it rather than passing it through; use the SDK's own output, or isHashedUserId to check first.

Hand-built OpenTelemetry spans

Hashing is applied at the SDK's own entry points. A span you build directly on an OTel tracer doesn't pass through them, so stamping user.id or session.id onto it yourself would emit a different value than the rest of your spans carry — the same person would appear as two users.

Apply the same processing explicitly:

import { processUserId, processSessionId } from "@runotis/sdk";

const span = tracer.startSpan("checkout.complete", { startTime });
span.setAttribute("user.id", processUserId(userId));
span.setAttribute("session.id", processSessionId(sessionId));

Both are idempotent, return undefined for undefined, and leave the value unchanged when identifier hashing is disabled — so they're safe to apply unconditionally. processArtifactId and the processUserIdAsync / processSessionIdAsync variants are also exported; prefer the async ones in browsers where the call site can await.

processArtifactId additionally trims and lowercases the id before hashing, so Deck_A1 and deck_a1 are the same artifact. It does this whether or not hashing is enabled, since the value is hashed byte-for-byte on ingest either way. User and session ids are not case-folded — those are treated as opaque, and Alice and alice remain different users.

Anonymous users

The browser SDK generates anonymous user IDs for unauthenticated sessions, prefixed anon_. These are hashed like any other user ID, so the value that reaches Otis carries the usr_v1_ prefix. When an anonymous visitor later signs in, identifyUser switches later activity to the signed-in ID. Otis doesn't merge the earlier anonymous activity into that user.

Server-side hashing layer

After values leave the SDK, a second hashing pass is applied in the ingest collector using a server-side secret. This is transparent to your application (you never see the second hash). You don't need to configure this layer; it happens automatically.

Why double-hash?

A single HMAC with a known secret is vulnerable to rainbow-table attacks on the space of likely raw identifiers: an attacker with the hashed values plus a guess at the input domain (email addresses, sequential internal IDs) can precompute hashes and match them against analytics storage. Double-hashing defeats this. Reversing a stored hash requires both the client-side seed and the server-side ingest secret, and neither is kept in analytics storage.

Who holds the client-side seed

By default Otis does. So that a new key can keep your users' identities, Otis records each environment's seed alongside the server-side secret, in its configuration store. The double hash therefore protects identifiers against a leak of analytics storage, and against anyone who holds only one of the two secrets. It does not protect them against a breach of Otis's configuration store together with analytics storage.

If you need the client-side hash to be a boundary Otis itself cannot cross, set identifierHashKey to a secret you generate and hold, and keep your API key server-side. Otis then never has the first of the two.

A key created before identity seeds, and not yet replaced, is as it was: Otis stores only a SHA-256 of it and cannot derive its seed.

Server-side keys only

A seed is only secret where the key that carries it is a server-side secret — a backend SDK, a serverless function, an edge worker. In a browser bundle the API key, or an identifierHashKey, is served to every visitor, and anyone can read the client-side seed from it.

For browser apps, treat client-side hashing as data minimization rather than as a second security boundary: raw emails and internal user IDs never travel over the wire, never enter request logs or proxy logs, and are never held in memory by Otis. The pseudonymization that resists reversal is the server-side pass, keyed by a per-project secret Otis never transmits to clients.

If you need identifiers that stay opaque to Otis in a browser context, hash them in your own backend before passing them to the SDK. Already-hashed values pass through unchanged.

Memories and uploaded documents

Redaction and identifier hashing apply to the telemetry your app sends. What your team tells Otis in a chat, and the files it uploads, are stored as provided. Keep personal data out of both. Memories and knowledge base describes who can see them.

Data deletion and export

To delete a project, to delete one user's data, or to export your data, contact Otis.

For production use, we recommend:

  • Leave client-side PII redaction on (the default).
  • Leave server-side PII redaction on (the default).
  • Leave identifier hashing enabled (the default).
  • Use on error: passthrough for analytics environments where data loss is costly (the default).
  • Use on error: drop for regulated environments where PII exposure is costlier than missing spans.

On this page