LP Agency

AI-automatic depersonalization system

documents for B2B

About the project

Type:
AI-MVP / Boxed B2B product
Direction:
Data Security / Business Automation
Branch:
Information security
Format:
Stand-alone on-premise solution
Link:
on request

The project was created as an independent B2B product for companies that work with a large number of documents and personal or confidential data.

The main idea is to automate the process of depersonalizing documents before transferring, storing, or using them in other business processes.

Instead of the employee manually searching the document for full name, document numbers, INN, names of organizations and other sensitive data, the system independently analyzes the contents, finds the necessary information and replaces it with special markers.

At the same time, the original data is not lost: the system retains the ability to restore them if appropriate access is available.

The product was designed as a boxed solution for installation inside the customer’s infrastructure. This was crucial for companies that cannot transfer documents to external cloud services.

tikcet2

Task

Create a system that automatically:

  • finds sensitive data in documents
  • depersonalizes them
  • assigns a unique code to the document
  • saves the protected version
  • allows you to restore data if necessary.

Context:

  • 438,000+ companies in the Russian Federation are personal data operators
  • most of them store documents on servers without full depersonalization.
  • Manual processing takes 18 times longer.
  • The main source of leaks is the human factor

The goal is to reduce the risks of data leakage through automated depersonalization.

The key difficulty

The main challenges:

  • Detection of sensitive data in arbitrary formats
  • Working with different languages, symbols, numbers, QR codes
  • Support for all Microsoft Office formats and analogues
  • Image Processing
  • Work in completely offline mode (without internet)
  • The possibility of secure data recovery

It’s not just text masking, it’s contextual document analysis.

Decision

The AI core

Two prototypes have been developed:

  • Depersonalization of text documents
  • Depersonalization of data in images

System:

  • analyzes the document
  • detects sensitive fields
  • replaces them with markers
  • assigns a unique identifier
  • saves a protected copy

The processing speed is up to 3 seconds per document.

Flexible setup

The company can:

  • choose data types to hide
  • configure depersonalization rules
  • choose the formats of the processed files

You can hide it:

  • FCs
  • document numbers
  • INN
  • QR codes
  • Latin alphabet
  • numbers
  • names of organizations

Architecture

  • A fully autonomous solution
  • Does not require an internet connection
  • On-premise installation
  • Digital signature for data recovery
  • Protection mechanisms against unauthorized photographing

tikcet3

Business value

Reducing leakage risks

Even if the server is hacked, the attacker receives anonymized data.

Saving time

If an employee manually spends 15 minutes depersonalizing a document, the system does it in 3 seconds.

15 minutes / 3 seconds = 300x faster
(conservatively— 18 times faster according to the pilot)

Compliance with regulatory requirements

The product reduces the risks of fines and reputational losses.

The monetization model

  • The boxed product
  • B2B segment
  • The price depends on the number of documents processed
  • Scaled by licenses

Scaling potential

The product can be developed sideways:

  • DLP integrations
  • API for enterprise systems
  • SaaS versions
  • AI models for risk forecasting
  • Integration with electronic document management

tikcet1

What is it really

It’s not just a depersonalization tool.

This is the security layer between the company’s server and the human factor.

If the server is hacked, the data is already protected.

Results

As part of the project, a working MVP of a boxed B2B product was created, rather than just a separate AI module.

At the exit, we received:

  • two directions of automatic depersonalization — documents and images;
  • context-based processing of sensitive data;
  • the ability to work with various types of personal and identifying data;
  • document processing speed up to 3 seconds;
  • autonomous architecture without transferring documents to external cloud services;
  • the mechanism for restoring the original data;
  • flexible configuration of depersonalization rules;
  • on-premise installation format for corporate infrastructure.

Thus, the technology was prepared not only as a demonstration of AI capabilities, but also as a basis for further development of a boxed product for the corporate market.