Privacy
Think You're Anonymous Once Your Name Is Gone?
mlj.solutions

mlj.solutions Smazané jméno nestačí. AI si ho může dopočítat. LLM agent umí propojit slabé stopy z více zdrojů a z anonymních dat rekonstruovat identitu. Pro firmy s AI přístupem k interním datům je to zásadní výzva. ❓Máte nastavenou privacy-by-design strategii pro vaše AI systémy?
View on Instagrammlj.solutions

mlj.solutions Smazali jste jméno. Data ještě anonymní nejsou. LLM agent dokáže spojovat slabé indicie - město, práci, zálibu, styl psaní, fragment příběhu - s veřejnými informacemi. A z něčeho, co samo o sobě nevypadá osobně, rekonstruovat reálnou identitu. To mění pravidla pro každou firmu, která dává AI přístup k dokumentům, CRM nebo interním znalostem. Nestačí řešit jen to, co je v jednom souboru. Musíte řešit, co jde odvodit kombinací více zdrojů. 👉 Jak řešíte privacy u interních AI systémů? Video ve čtvrtek.
View on InstagramWhat the Study Is About
Anonymized data has long been considered safe. This study shows that LLM agents can piece together a specific identity from scattered, individually non-identifying clues — without any specialized engineering.
79.2%
of anonymous profiles successfully de-anonymized
Agents reconstructed 792 out of 1,000 identities in the Netflix dataset, while the classic method reached only 56%.
98%
success rate on an explicit request
When directly asked to re-identify, Claude 4.5 achieved over 98% success across all types of fingerprint clues.
10
identities uncovered from AOL search logs
The agent linked anonymous queries to public sources and exposed specific individuals, including highly sensitive health and personal details.
What This Means
- Anonymization alone isn't enough, removing a name and direct identifiers doesn't help if contextual clues remain.
- De-anonymization doesn't have to be intentional, an agent can reconstruct an identity as a side effect of an ordinary analytical task.
- Safety guardrails fail, explicit requests for re-identification aren't consistently refused by models.
- Multiple weak clues add up, no single piece of information identifies someone, but their combination does.
- Protection requires active design, a privacy-aware system prompt reduces risk, but at the cost of the model's usefulness.
Our Take
The biggest security risk isn't an attacker. It's blind trust — the belief that a tool can be relied on 100%.
Michal Dobrovolný
Zdroj
Ko, M., Jeong, J., Thakur, S. S., Kim, G., & Jia, R. (2026). From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents. arXiv:2603.18382. https://arxiv.org/abs/2603.18382
Keep reading
Don't miss these posts
Want to implement AI in your company?
We'll go through your needs together and propose an approach that delivers measurable results.
Write to us