What This Demonstrates
A document full of personal data can be processed by a remote model API without any of that personal data leaving your machine.
The Three Steps
1. Detect + Anonymize
Emails, phones and URLs are found with deterministic patterns; names and locations with NER (Microsoft Presidio, running locally). Each value is replaced with a token like <<EMAIL_01>> and the token→value map is kept in memory here.
2. Send
The tokenized text goes to the configured model endpoint — by default a local Ollama. A tripwire re-scans the outbound payload first and blocks the send if any real value remains. The map is never part of the request, so it cannot be leaked or prompt-injected out. It simply isn't there.
3. Rehydrate
Tokens in the response are swapped back for real values, locally. Any token the model altered is surfaced as an error, never silently dropped.
Honest Limitations
Detection is probabilistic. A missed name is sent in the clear, which is why recall must be measured, not assumed.
Tokenization also does not solve re-identification from context: "sole female VP at a 40-person Reykjavík fintech" identifies someone with every name removed.
Runtime Log
Every line below passes the same redaction filter as outbound payloads, so no real value can appear here.How Log Redaction Works
Values shorter than 4 characters (e.g. a state code) are
deliberately not redacted. Corpus filenames are registered as
secrets at startup, so even Loading <file>.pdf
shows [REDACTED].
Pick a document and press Run.
Nothing yet.
Nothing yet.