Deterministic anonymization: join keys without storing PII — Veramask

The first thing anonymization breaks is correlation. The moment every call replaces a real email with a fresh fake one, your analytics stops grouping, your logs stop tracing, and your dedupe stops deduplicating — because the same person now looks like a different person on every line. Random anonymization is safe and useless in the same breath: it removed the identity you were trying to study along with the identity you were trying to hide.

The fix is a salt. When anonymization is deterministic — the same input plus the same salt always produces the same output — a masked value becomes a stable handle for the record it came from. You get the join key of a tokenization vault without operating, securing, or breaching a vault. This article walks through the join-key problem, how salted determinism works without any server-side state, worked examples for the hash and substitute strategies, how to govern the salt, and when determinism is the wrong choice.

Deterministic anonymization: with a salt the same input produces the same token on every call; without it, every call produces a different token Two panels. Left panel: with a consistency salt, the input jane.doe@example.com produces the identical token 8f223fc1a0b9 on call one, call two, and call three — a stable join key. Right panel: without a salt, the same input produces a different token on each call, destroying identity. The salt lives in the caller's configuration, never on the server. Same input, three calls — with and without a salt with a salt input (every call) jane.doe@example.com call 1 8f223fc1a0b9 call 2 8f223fc1a0b9 call 3 8f223fc1a0b9 the same token every time — a stable join key join logs, group users, trace across services without a salt input (every call) jane.doe@example.com call 1 8f223fc1a0b9 call 2 c41d9e7723ab call 3 5ab02e8847f1 a fresh token per call — identity is destroyed no joins, no group-bys, no tracing

The salt lives in your configuration. The service recomputes the same function and keeps nothing.

The direct answer

How do you keep analytics working after anonymization? Make the anonymization deterministic with a salt you keep. In the Veramask API, that means passing a consistency_salt in settings: from then on, the hash and substitute strategies produce the same output for the same input on every call — in every service, every environment, every CI run. The token a log line carries is the same token the same user produced in every other log line, so you can group, join, and trace on it exactly as you would on the original value. The reproducibility lives in the function and in your salt, not in anything the service remembers — the same request yields the same result because the transformation is a pure computation, not a lookup against stored data.

What breaks when anonymization is random

What you were doing What happens without determinism
GROUP BY user in analytics One real user fragments into N synthetic users per request — counts inflate, averages smear
Sessionization and funnels Steps can't be linked to one subject, so funnels leak at every stage
Cross-service tracing Service A's log and service B's event carry different fake values for the same person
Duplicate detection Every copy of a record looks unique; dedupe logic never fires
Joining logs to a warehouse No shared key exists on either side
"Top users by activity" Rankings become noise; heavy users don't concentrate on one token

The underlying problem is that the anonymized value was carrying the identity. Random substitution is supposed to destroy that identity — that is its privacy job. The mistake is applying it where the identity was the analytical key. Determinism restores the key while still never emitting the original value: the token is a handle to the record, not the record.

How salted determinism works

There is no magic and no memory. When you provide a consistency_salt, the service mixes it into the transformation:

  • hash computes a salted SHA-256 of the detected value. Same input + same salt → the same 64-hex-character token, every time.
  • substitute selects the synthetic value as a deterministic function of the input and the salt. Same input + same salt → the same fake name, email, or phone number, every time.

The salt is the only cross-request link, and it is yours. The service holds your payload in volatile memory for the duration of the request, applies the function, returns the result, and keeps nothing — no value cache, no token store, no mapping table. Two identical requests produce identical outputs because the computation is identical, not because anything was remembered. That is the difference between deterministic anonymization and a tokenization service: the vault has been replaced by arithmetic. The architectural contract behind that claim is described in why stateless anonymization reduces exposure.

The rules, from the API contract:

Setting Rule
settings.consistency_salt At least 16 bytes (UTF-8) when provided. Generate it randomly — 32 hex characters is a comfortable size.
Strategies affected hash and substitute. redact and replace destroy the value's identity by design; mask preserves part of the original and is not a privacy boundary.
Availability Plan-gated. Passing a salt from a plan that doesn't include it returns 403 Subscription Limit — it is never silently ignored. Check your plan on the Subscription page.

Example: a join key across log lines (hash)

The classic use case: sanitize logs for PII while keeping them traceable. Hash the entities that identify a user, with a stable salt, and the token becomes the correlation key:

curl -X POST https://api.veramask.com/v1/anonymizeText \
  -H "x-api-key: $VERAMASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "data": "User Jane Doe (jane.doe@example.com) failed checkout from 203.0.113.42",
    "settings": {
      "locale": "en_US",
      "confidence_threshold": 0.8,
      "consistency_salt": "0123456789abcdef"
    },
    "overrides": {
      "EMAIL_ADDRESS": { "type": "hash" },
      "IP_ADDRESS": { "type": "hash" }
    }
  }'
{
  "masked_data": "User Jane Doe (9f2c1a...) failed checkout from 8f223fc..."
}

Every log line that mentions jane.doe@example.com now carries the same 9f2c1a... token. Grep for the token and you have the complete trail of that user across your log estate — without the email existing in any of those lines. Put the same salt in service B's sanitizer configuration and service B computes the same token for the same email, so the two systems correlate with zero coordination: no shared token service in the request path, no lookup calls, no state to keep warm. The join key is derived independently on both sides.

Example: a stable pseudonym per subject (substitute)

Hashes are ugly in datasets humans read. substitute with a salt gives you the same determinism in realistic shape — the same subject maps to the same synthetic identity across every dataset processed with that salt:

curl -X POST https://api.veramask.com/v1/anonymizeText \
  -H "x-api-key: $VERAMASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "data": "Alice Smith can be reached at alice@example.com",
    "settings": {
      "locale": "en_US",
      "consistency_salt": "0123456789abcdef"
    },
    "overrides": {
      "PERSON": { "type": "substitute", "attribute": "name" },
      "EMAIL_ADDRESS": { "type": "substitute", "attribute": "email" }
    }
  }'
{
  "masked_data": "Christopher Taylor can be reached at pespinoza@example.com"
}

Every occurrence of Alice Smith in the salted corpus becomes Christopher Taylor; every occurrence of her email becomes pespinoza@example.com. A staging database seeded this way keeps working — group-bys return stable per-user rows, foreign keys stay consistent between tables processed with the same salt, and the data still looks like data. One caveat carried over from the strategy reference: substitution does not preserve relationships between fields. The fake name and the fake email are each individually stable, but they do not describe the same fictional person — if you need referential integrity across entity types, use hash as the join key and substitute for the human-readable surfaces.

Salt governance: the salt is a credential

Determinism moves the secret from the database to the salt. Treat it accordingly:

  • Generate it randomly and store it in a secret manager — the same place your API credentials live, never in the repository, never in code.
  • Never log it next to the tokens it produces. A salt published beside its hashes converts every pseudonym back into a verifiable guess target.
  • Scope it per purpose. One salt for log sanitization, a different one for the dataset you share with a third party. A token from one scope cannot be correlated with the other, and a leak in one scope doesn't compromise the other.
  • Rotate it like a credential — knowing the cost. A new salt re-baselines every token: all historical tokens stop matching new ones. Version the epoch (a token_epoch column, or a prefix on the salt itself) so consumers can tell which generation a token belongs to, and keep the old salt only as long as the migration needs it. Then destroy it.
  • Deleting the salt is the off switch. With the salt gone, tokens can no longer be verified or re-derived against candidate values. The pseudonymization link you created is severed — which is exactly what you want when the analytical purpose expires.

When not to use determinism

Determinism is a default for pipelines, not a law of nature. Skip the salt when:

  • The population is small. A stable token per subject preserves cardinality and frequency. In a five-user dataset, the token is the user — context alone re-identifies them. If the same fake value recurring across records would itself be a leak, you want per-request randomness, not consistency.
  • The field is low-cardinality. Fifty states map to fifty tokens; a frequency analysis matches them in minutes. Deterministic substitution of LOCATION buys you nothing analytically and confirms structure to an attacker.
  • The job is one-shot. A single log line that will never be joined to anything gains nothing from a salt — and every salt you manage is a credential with a rotation cost.
  • Freshness is the point. Demo screenshots, red-teaming fixtures, and load-test payloads often want different fake data on every run. Without a salt, that is exactly what you get.

A token vault, minus the vault

The enterprise alternative to this pattern is a tokenization vault: an encrypted store mapping originals to tokens, with a detokenization API for authorized callers. Vaults work — they are also a concentrated database of originals that you must run, secure, back up, and audit forever. Salted determinism gets you the same correlation property with a different architecture:

Tokenization vault Salted determinism
What it is A database of original ↔ token mappings A secret plus a pure function
Server-side store Yes — the vault holds every original None — nothing to store, nothing to breach
Reversible Yes, by design (detokenization) hash: not decryptable — with the salt you can verify a guess, never recover the value. substitute: not derivable at all
Operational cost Run, secure, back up, audit, and scale the vault Keep one secret in the secret manager
Breach impact The vault is a single point of total PII compromise A leaked salt is a rotated credential; there is no data store behind it
GDPR posture Pseudonymisation with added attack surface Pseudonymisation with subtractive storage

If you need the original back later — for support workflows that must re-identify a customer — you need a vault or encryption, not anonymization, and you should think carefully about who holds that key. If you need correlation without recovery, the salt is the whole answer.

FAQ

Is a salted hash anonymous data?

No. It is pseudonymisation. Because the salt plus the token lets you verify candidate values, GDPR treats the output as personal data — the regulation recognizes exactly this technique, under Article 32(1)(a), as a security measure rather than an exemption. That is the right way to use it: a strong improvement over plaintext with the correlation you need. If your goal is true anonymization in the strict sense — data that exits GDPR scope — the tools are deletion, aggregation, and noise, not hashing. The strategy reference covers which transformations get you there.

Does the service remember my salt or my values?

No. The salt is used per request, in volatile memory, and never persisted. Identical requests yield identical outputs because the transformation is a deterministic function of (input, salt) — nothing is cached, looked up, or learned. The reproducibility lives in the function and in the salt you keep; see the stateless contract for what the service guarantees it does not hold.

What happens when I rotate the salt?

Every token changes. Treat rotation as a re-baseline of your analytics: new tokens stop matching historical ones, so either version the token generation and keep both epochs queryable, or accept the cliff and reprocess what you can. Keep the old salt only for the migration window, then destroy it — a retired salt has no analytical value and nonzero risk.

Can I use one salt for everything?

Technically yes; hygienically, scope per purpose. A single salt means one leak compromises every dataset's linkability, and tokens from your internal logs could be correlated with tokens in a dataset you shared externally. Separate salts per purpose are free — they're just more entries in the secret manager — and they cap the blast radius of any one disclosure.

Why do I get different outputs for the same input?

Three usual causes: no consistency_salt in settings (substitutions and unsalted hashes vary by request — that is documented behavior); a different salt between the two calls; or your plan doesn't include the salt feature, in which case the request fails with 403 Subscription Limit rather than silently ignoring it. Confirm the exact same salt string is being sent every time before anything else.

Connect with us