Browser Extension Anonymizes Sensitive Data Before Sending to AI Chatbots
A new browser extension has been developed to solve the growing problem of sensitive data leaking into AI chat services through everyday work tasks.
HR specialists, lawyers, accountants, and IT staff routinely paste resumes, contracts, payroll tables, and configuration files containing personal data, bank details, and passwords directly into AI prompts. Once sent, this information remains on third-party servers with no reliable way to delete it from logs or training datasets.
The extension intercepts messages and file uploads at the moment of sending, replaces detected sensitive values with unique pseudonyms such as [user_qhb6tm] or [card_348zsg], and forwards only the masked text to the AI. When the AI replies using the same placeholders, the extension instantly restores the original values on the user's screen while the server sees only the anonymized version.
Earlier approaches proved insufficient. Manual replacement quickly failed due to user fatigue. Simple regex scripts struggled with Russian grammar, false positives on invoice numbers, and inconsistent name formats. External anonymization services merely shifted the trust problem to another third party.
The final solution meets three strict requirements: fully automatic operation without user intervention, completely local processing with no internet required after installation, and seamless integration that does not change how people interact with their preferred AI chat interface.
Inside the engine, 33 data categories are defined declaratively with patterns, validators, and priorities. INN numbers are verified using FNS checksum algorithms, cards use the Luhn check, and passwords are filtered by entropy and stop-word lists. Russian names are normalized to nominative case so that declined forms receive the same pseudonym. The response parser tolerates case changes, spaces, markdown, and homoglyphs to ensure reliable decoding even when models alter labels.
File support includes more than 70 formats. Office documents are processed by creating anonymized copies that preserve formatting, tables, and headers. Scanned PDFs and photos are handled by a local offline OCR engine that runs in a single worker thread, keeping memory usage low and the system responsive.
A free version covering all text categories and custom rules is available in the Chrome Web Store. Advanced features for files, scans, and enterprise policy distribution are provided upon request.
Related articles
Amnezia VPN Survives Coordinated Russian Censorship Campaign Targeting AmneziaWG Protocol Fingerprints
Amnezia VPN has published a detailed post-mortem on the multi-wave blocking campaign conducted by Russian authorities against its Amnezia Free and Amnezia Premium services during June and July. The company describes a shift from simple protocol blocking to sophisticated fingerprinting of AmneziaWG traffic combined with infrastructure DDoS attacks and automated IP-subnet blacklisting. Engineers closed multiple detection vectors including zero-length UDP packets, fixed-size keepalive messages, handshake timing patterns, and nonce zero bytes. The incident forced accelerated migration to AmneziaWG 2.0, discontinuation of legacy client support, and development of AmneziaWG 3.0 while expanding VLESS infrastructure as a backup. Self-hosted users largely avoided direct protocol blocks but still faced subnet-level restrictions. The report highlights how Roskomnadzor now applies cumulative scoring across multiple traffic features rather than single definitive markers.
Data Masking: 8 Critical Questions Businesses and Developers Ask About Protecting Sensitive Data
Garda expert Dmitry Larin addresses common challenges in data masking during a recent webinar titled 'Data Masking: Battle of Opinions'. The discussion covers why masking remains essential even when encryption is deployed, how to preserve application functionality after anonymization, and the performance trade-offs of processing large databases such as 5 TB PostgreSQL instances. Different masking types including static, dynamic, selective, and streaming are explained with specific use cases for DevOps pipelines, external contractors, and BI systems. The article also examines why machine learning alone is insufficient for discovering personal data and why custom scripts fail at scale across heterogeneous environments like PostgreSQL and Oracle. Practical recommendations include combining masking with encryption, using deterministic transformations for deduplication, and separating replication from masking tasks to avoid production impact.
MAX Desktop Client Tested for VPN Detection on Windows, No Tracking Signs Found
A Habra user named Slava_B conducted an experiment on September 8, 2026, to determine whether the MAX desktop client on Windows could detect or route traffic through a VPN configured at the router level. The setup used a Keenetic router that directed Russian resources directly while sending other connections via an OpenConnect tunnel to a European VPS, with no VPN client or virtual adapter present in Windows itself. Monitoring tools including Process Monitor, Wireshark, TCPView, and tcpdump revealed that MAX.exe and MAX-service.exe processes communicate locally and connect to MAX/ONEME infrastructure along with AppTracer services. The application repeatedly accessed MachineGuid, computer name, proxy settings, device IDs, and microphone/camera information, though these reads may support diagnostics and anti-fraud functions. No connections appeared on the VPN interface, and the client did not attempt to reach IP-checking services, Telegram, or WhatsApp. The researcher noted that TLS traffic was not decrypted, so actual transmission of identifiers could not be confirmed, and results apply only to this router-based configuration.
PII-Guard: Open-Source Detector for Personal Data in Russian Text
Andrey Ivanov, an NLP researcher at red_mad_robot, has released PII-Guard, an open-source system that detects and masks personal data in Russian text before it reaches language models. The tool combines rule-based checks with a fine-tuned ruBert-base NER model to handle names, addresses, phones, passports, INN, SNILS, bank cards and other entities. It replaces detected PII with structured XML-like tags that preserve grammatical information such as gender and entity ID, allowing models to generate coherent responses that are later restored with real values. The hybrid pipeline first applies normalization, pattern matching, Luhn and weighted checksum validation, and context windows with positive and negative keywords, then merges results with model predictions via an arbitration module. Evaluation on four public datasets, including Hivetrace, alexen2 and alrosait, shows PII-Guard outperforming other open solutions on both strict span matching and type-overlap micro-F1 metrics. The project, including datasets and code, is available on GitHub and aims to reduce leakage risks while maintaining downstream model utility.