The first thing anonymization breaks is correlation. The moment every call replaces a real email with a fresh fake one, your analytics stops grouping, your logs stop tracing, and your dedupe stops deduplicating — because the same person now looks like a different person on every line. Random anonymization is safe and useless in the same breath: it removed the identity you were trying to study along with the identity you were trying to hide.
The fix is a salt. When anonymization is deterministic — the same input plus
the same salt always produces the same output — a masked value becomes a
stable handle for the record it came from. You get the join key of a
tokenization vault without operating, securing, or breaching a vault. This
article walks through the join-key problem, how salted determinism works
without any server-side state, worked examples for the hash and
substitute strategies, how to govern the salt, and when determinism is the
wrong choice.
The direct answer
How do you keep analytics working after anonymization? Make the
anonymization deterministic with a salt you keep. In the Veramask API, that
means passing a consistency_salt in settings: from then on, the hash
and substitute strategies produce the same output for the same input on
every call — in every service, every environment, every CI run. The token a
log line carries is the same token the same user produced in every other log
line, so you can group, join, and trace on it exactly as you would on the
original value. The reproducibility lives in the function and in your salt,
not in anything the service remembers — the same request yields the same
result because the transformation is a pure computation, not a lookup against
stored data.
What breaks when anonymization is random
| What you were doing | What happens without determinism |
|---|---|
GROUP BY user in analytics |
One real user fragments into N synthetic users per request — counts inflate, averages smear |
| Sessionization and funnels | Steps can't be linked to one subject, so funnels leak at every stage |
| Cross-service tracing | Service A's log and service B's event carry different fake values for the same person |
| Duplicate detection | Every copy of a record looks unique; dedupe logic never fires |
| Joining logs to a warehouse | No shared key exists on either side |
| "Top users by activity" | Rankings become noise; heavy users don't concentrate on one token |
The underlying problem is that the anonymized value was carrying the identity. Random substitution is supposed to destroy that identity — that is its privacy job. The mistake is applying it where the identity was the analytical key. Determinism restores the key while still never emitting the original value: the token is a handle to the record, not the record.
How salted determinism works
There is no magic and no memory. When you provide a consistency_salt, the
service mixes it into the transformation:
hashcomputes a salted SHA-256 of the detected value. Same input + same salt → the same 64-hex-character token, every time.substituteselects the synthetic value as a deterministic function of the input and the salt. Same input + same salt → the same fake name, email, or phone number, every time.
The salt is the only cross-request link, and it is yours. The service holds your payload in volatile memory for the duration of the request, applies the function, returns the result, and keeps nothing — no value cache, no token store, no mapping table. Two identical requests produce identical outputs because the computation is identical, not because anything was remembered. That is the difference between deterministic anonymization and a tokenization service: the vault has been replaced by arithmetic. The architectural contract behind that claim is described in why stateless anonymization reduces exposure.
The rules, from the API contract:
| Setting | Rule |
|---|---|
settings.consistency_salt |
At least 16 bytes (UTF-8) when provided. Generate it randomly — 32 hex characters is a comfortable size. |
| Strategies affected | hash and substitute. redact and replace destroy the value's identity by design; mask preserves part of the original and is not a privacy boundary. |
| Availability | Plan-gated. Passing a salt from a plan that doesn't include it returns 403 Subscription Limit — it is never silently ignored. Check your plan on the Subscription page. |
Example: a join key across log lines (hash)
The classic use case: sanitize logs for PII while keeping them traceable. Hash the entities that identify a user, with a stable salt, and the token becomes the correlation key:
curl -X POST https://api.veramask.com/v1/anonymizeText \
-H "x-api-key: $VERAMASK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"data": "User Jane Doe (jane.doe@example.com) failed checkout from 203.0.113.42",
"settings": {
"locale": "en_US",
"confidence_threshold": 0.8,
"consistency_salt": "0123456789abcdef"
},
"overrides": {
"EMAIL_ADDRESS": { "type": "hash" },
"IP_ADDRESS": { "type": "hash" }
}
}'
{
"masked_data": "User Jane Doe (9f2c1a...) failed checkout from 8f223fc..."
}
Every log line that mentions jane.doe@example.com now carries the same
9f2c1a... token. Grep for the token and you have the complete trail of that
user across your log estate — without the email existing in any of those
lines. Put the same salt in service B's sanitizer configuration and service B
computes the same token for the same email, so the two systems correlate with
zero coordination: no shared token service in the request path, no lookup
calls, no state to keep warm. The join key is derived independently on both
sides.
Example: a stable pseudonym per subject (substitute)
Hashes are ugly in datasets humans read. substitute with a salt gives you
the same determinism in realistic shape — the same subject maps to the same
synthetic identity across every dataset processed with that salt:
curl -X POST https://api.veramask.com/v1/anonymizeText \
-H "x-api-key: $VERAMASK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"data": "Alice Smith can be reached at alice@example.com",
"settings": {
"locale": "en_US",
"consistency_salt": "0123456789abcdef"
},
"overrides": {
"PERSON": { "type": "substitute", "attribute": "name" },
"EMAIL_ADDRESS": { "type": "substitute", "attribute": "email" }
}
}'
{
"masked_data": "Christopher Taylor can be reached at pespinoza@example.com"
}
Every occurrence of Alice Smith in the salted corpus becomes
Christopher Taylor; every occurrence of her email becomes
pespinoza@example.com. A staging database seeded this way keeps working —
group-bys return stable per-user rows, foreign keys stay consistent between
tables processed with the same salt, and the data still looks like data.
One caveat carried over from the
strategy reference:
substitution does not preserve relationships between fields. The fake name
and the fake email are each individually stable, but they do not describe the
same fictional person — if you need referential integrity across entity
types, use hash as the join key and substitute for the human-readable
surfaces.
Salt governance: the salt is a credential
Determinism moves the secret from the database to the salt. Treat it accordingly:
- Generate it randomly and store it in a secret manager — the same place your API credentials live, never in the repository, never in code.
- Never log it next to the tokens it produces. A salt published beside its hashes converts every pseudonym back into a verifiable guess target.
- Scope it per purpose. One salt for log sanitization, a different one for the dataset you share with a third party. A token from one scope cannot be correlated with the other, and a leak in one scope doesn't compromise the other.
- Rotate it like a credential — knowing the cost. A new salt re-baselines
every token: all historical tokens stop matching new ones. Version the
epoch (a
token_epochcolumn, or a prefix on the salt itself) so consumers can tell which generation a token belongs to, and keep the old salt only as long as the migration needs it. Then destroy it. - Deleting the salt is the off switch. With the salt gone, tokens can no longer be verified or re-derived against candidate values. The pseudonymization link you created is severed — which is exactly what you want when the analytical purpose expires.
When not to use determinism
Determinism is a default for pipelines, not a law of nature. Skip the salt when:
- The population is small. A stable token per subject preserves cardinality and frequency. In a five-user dataset, the token is the user — context alone re-identifies them. If the same fake value recurring across records would itself be a leak, you want per-request randomness, not consistency.
- The field is low-cardinality. Fifty states map to fifty tokens; a
frequency analysis matches them in minutes. Deterministic substitution of
LOCATIONbuys you nothing analytically and confirms structure to an attacker. - The job is one-shot. A single log line that will never be joined to anything gains nothing from a salt — and every salt you manage is a credential with a rotation cost.
- Freshness is the point. Demo screenshots, red-teaming fixtures, and load-test payloads often want different fake data on every run. Without a salt, that is exactly what you get.
A token vault, minus the vault
The enterprise alternative to this pattern is a tokenization vault: an encrypted store mapping originals to tokens, with a detokenization API for authorized callers. Vaults work — they are also a concentrated database of originals that you must run, secure, back up, and audit forever. Salted determinism gets you the same correlation property with a different architecture:
| Tokenization vault | Salted determinism | |
|---|---|---|
| What it is | A database of original ↔ token mappings | A secret plus a pure function |
| Server-side store | Yes — the vault holds every original | None — nothing to store, nothing to breach |
| Reversible | Yes, by design (detokenization) | hash: not decryptable — with the salt you can verify a guess, never recover the value. substitute: not derivable at all |
| Operational cost | Run, secure, back up, audit, and scale the vault | Keep one secret in the secret manager |
| Breach impact | The vault is a single point of total PII compromise | A leaked salt is a rotated credential; there is no data store behind it |
| GDPR posture | Pseudonymisation with added attack surface | Pseudonymisation with subtractive storage |
If you need the original back later — for support workflows that must re-identify a customer — you need a vault or encryption, not anonymization, and you should think carefully about who holds that key. If you need correlation without recovery, the salt is the whole answer.
FAQ
Is a salted hash anonymous data?
No. It is pseudonymisation. Because the salt plus the token lets you verify candidate values, GDPR treats the output as personal data — the regulation recognizes exactly this technique, under Article 32(1)(a), as a security measure rather than an exemption. That is the right way to use it: a strong improvement over plaintext with the correlation you need. If your goal is true anonymization in the strict sense — data that exits GDPR scope — the tools are deletion, aggregation, and noise, not hashing. The strategy reference covers which transformations get you there.
Does the service remember my salt or my values?
No. The salt is used per request, in volatile memory, and never persisted. Identical requests yield identical outputs because the transformation is a deterministic function of (input, salt) — nothing is cached, looked up, or learned. The reproducibility lives in the function and in the salt you keep; see the stateless contract for what the service guarantees it does not hold.
What happens when I rotate the salt?
Every token changes. Treat rotation as a re-baseline of your analytics: new tokens stop matching historical ones, so either version the token generation and keep both epochs queryable, or accept the cliff and reprocess what you can. Keep the old salt only for the migration window, then destroy it — a retired salt has no analytical value and nonzero risk.
Can I use one salt for everything?
Technically yes; hygienically, scope per purpose. A single salt means one leak compromises every dataset's linkability, and tokens from your internal logs could be correlated with tokens in a dataset you shared externally. Separate salts per purpose are free — they're just more entries in the secret manager — and they cap the blast radius of any one disclosure.
Why do I get different outputs for the same input?
Three usual causes: no consistency_salt in settings (substitutions and
unsalted hashes vary by request — that is documented behavior); a different
salt between the two calls; or your plan doesn't include the salt feature, in
which case the request fails with 403 Subscription Limit rather than
silently ignoring it. Confirm the exact same salt string is being sent every
time before anything else.
Connect with us
- Follow @veramaskapi on X
- Try the API at veramask.com
- Read the transformation strategies and API specifications docs
