Engineer Details Six Weeks Spent Training and Testing Signature Redaction Models for Closed-Loop Document Anonymization
A developer working on automated document sanitization spent six weeks attempting to reliably locate and redact handwritten signatures inside PDFs that must leave a closed network. The goal was to remove surnames, telephone numbers, addresses and signatures while preserving document usability, all on a single air-gapped workstation equipped with one consumer GPU.
Documents arrive in two forms. Text-layer PDFs allow direct removal of recognized words. Scanned pages require the system to detect text via Tesseract, then draw black rectangles over the corresponding image regions. Four parallel detection methods—dictionary lookup, regular expressions, natasha NER and an ink gate—successfully handled printed names but found zero signatures because signatures contain no machine-readable text.
Initial experiments with a color-saturation gate worked on blue-ink signatures yet failed completely on black-ink signatures and on any grayscale scan. An alternative heuristic that painted any dark region ignored by OCR erased entire table columns when text was printed in small fonts, violating the rule that only recognized text may be redacted.
Three off-the-shelf object detectors were benchmarked on 1,254 pages. YOLOS produced excessive false positives on engineering drawings and was discarded. A seal detector returned high-confidence false positives on documents containing no seals. A custom YOLO11s model trained on 600 synthetically augmented signature images plus 40 real pages reached 88 % coverage on isolated strokes but dropped sharply when adjacent signatures merged into large blobs.
Two rotation-related bugs repeatedly caused missed redactions. PDF /Rotate attributes produced coordinate mismatches between detection and drawing layers. Pages scanned sideways defeated Tesseract’s orientation detection when run with --psm 0, leaving the entire page unprocessed until OSD was re-enabled for low-confidence cases.
After switching the vision model from proactive detection to post-redaction verification, expensive GPU calls fell from roughly 50 pages per batch to 18. On the final corpus of 492 pages only 3.7 % of documents required the vision stage, all of them empty forms with vertical table headers.
Image-preprocessing tests showed that the classical OpenCV pipeline (deskew, illumination normalization, denoising, 2× upscaling) raised OCR to 92/103 pages while RealESRGAN reduced the same metric to 76/103 and increased runtime dramatically. Preprocessing is therefore applied only when the first-pass score falls below 70.
Comparison of twelve OCR engines on a two-page test document revealed that PaddleOCR detection combined with Tesseract recognition produced 16 correctly grouped name cells in 3.7 seconds. Full replacement by any single vision-language model either missed names entirely or required 12–20 seconds per page.
Redaction implementation errors were discovered only after the author added a verification step that extracts text from beneath every black rectangle. Using PyMuPDF draw_rect left recoverable text; switching to apply_redactions eliminated all leaks across 3,029 redactions in production files. White rectangles drawn for stamps and signatures occasionally overwrote already-redacted black areas, reopening rows of names—an issue caught only by the same under-rectangle text check.
The final pipeline therefore combines rule-based text removal, a lightweight YOLO signature detector, selective classical preprocessing, Tesseract augmented by PaddleOCR cell detection, and mandatory post-redaction text verification, achieving reliable signature removal inside the required closed environment.
Related articles
Amnezia VPN Survives Coordinated Russian Censorship Campaign Targeting AmneziaWG Protocol Fingerprints
Amnezia VPN has published a detailed post-mortem on the multi-wave blocking campaign conducted by Russian authorities against its Amnezia Free and Amnezia Premium services during June and July. The company describes a shift from simple protocol blocking to sophisticated fingerprinting of AmneziaWG traffic combined with infrastructure DDoS attacks and automated IP-subnet blacklisting. Engineers closed multiple detection vectors including zero-length UDP packets, fixed-size keepalive messages, handshake timing patterns, and nonce zero bytes. The incident forced accelerated migration to AmneziaWG 2.0, discontinuation of legacy client support, and development of AmneziaWG 3.0 while expanding VLESS infrastructure as a backup. Self-hosted users largely avoided direct protocol blocks but still faced subnet-level restrictions. The report highlights how Roskomnadzor now applies cumulative scoring across multiple traffic features rather than single definitive markers.
Data Masking: 8 Critical Questions Businesses and Developers Ask About Protecting Sensitive Data
Garda expert Dmitry Larin addresses common challenges in data masking during a recent webinar titled 'Data Masking: Battle of Opinions'. The discussion covers why masking remains essential even when encryption is deployed, how to preserve application functionality after anonymization, and the performance trade-offs of processing large databases such as 5 TB PostgreSQL instances. Different masking types including static, dynamic, selective, and streaming are explained with specific use cases for DevOps pipelines, external contractors, and BI systems. The article also examines why machine learning alone is insufficient for discovering personal data and why custom scripts fail at scale across heterogeneous environments like PostgreSQL and Oracle. Practical recommendations include combining masking with encryption, using deterministic transformations for deduplication, and separating replication from masking tasks to avoid production impact.
MAX Desktop Client Tested for VPN Detection on Windows, No Tracking Signs Found
A Habra user named Slava_B conducted an experiment on September 8, 2026, to determine whether the MAX desktop client on Windows could detect or route traffic through a VPN configured at the router level. The setup used a Keenetic router that directed Russian resources directly while sending other connections via an OpenConnect tunnel to a European VPS, with no VPN client or virtual adapter present in Windows itself. Monitoring tools including Process Monitor, Wireshark, TCPView, and tcpdump revealed that MAX.exe and MAX-service.exe processes communicate locally and connect to MAX/ONEME infrastructure along with AppTracer services. The application repeatedly accessed MachineGuid, computer name, proxy settings, device IDs, and microphone/camera information, though these reads may support diagnostics and anti-fraud functions. No connections appeared on the VPN interface, and the client did not attempt to reach IP-checking services, Telegram, or WhatsApp. The researcher noted that TLS traffic was not decrypted, so actual transmission of identifiers could not be confirmed, and results apply only to this router-based configuration.
PII-Guard: Open-Source Detector for Personal Data in Russian Text
Andrey Ivanov, an NLP researcher at red_mad_robot, has released PII-Guard, an open-source system that detects and masks personal data in Russian text before it reaches language models. The tool combines rule-based checks with a fine-tuned ruBert-base NER model to handle names, addresses, phones, passports, INN, SNILS, bank cards and other entities. It replaces detected PII with structured XML-like tags that preserve grammatical information such as gender and entity ID, allowing models to generate coherent responses that are later restored with real values. The hybrid pipeline first applies normalization, pattern matching, Luhn and weighted checksum validation, and context windows with positive and negative keywords, then merges results with model predictions via an arbitration module. Evaluation on four public datasets, including Hivetrace, alexen2 and alrosait, shows PII-Guard outperforming other open solutions on both strict span matching and type-overlap micro-F1 metrics. The project, including datasets and code, is available on GitHub and aims to reduce leakage risks while maintaining downstream model utility.