Why is this needed? The accounting department is tired of manually entering all the data from passports / SNILS / medical books. Well, let's fix that.

At first glance, the task sounds simple: you upload a photo or scan of the passport, and you get letters from the surname to "issued by"—don't forget about the series and number.

Let's go!

So, OCR can return some text.

But in practice, it turned out that "just recognizing text" and "extracting normal fields from the passport" are two different tasks. And the second one turned out to be more complex.

The accounting department needs not just text, but a neat structure that can then be given to the UI for manual verification and sent to the API / 1C / CRM afterward.

First version: just run it through PaddleOCR

I chose PaddleOCR with Russian language support as the OCR engine.

At the start, the idea is as simple as possible:

  1. The user uploads a file.
  2. The service converts it to a normal format.
  3. PaddleOCR recognizes the text.
  4. The parser tries to extract the passport fields from the text.

I allowed basic formats for input: jpg, jpeg, png, pdf.

Even if the user uploads a regular image, I still normalize it to an internal PNG. This is convenient because the entire pipeline then works with one format.

And that's where the problems began.

First problem: OCR returns not a document, but a set of pieces

PaddleOCR doesn't return text nicely right away. It gives a set of found text blocks, i.e., the text itself, confidence score, block coordinates, bbox, and other stuff.

So the output doesn't look like this:

Last name: Ivanov
First name: Ivan

But rather a chaotic list of OCR boxes that need to be assembled back into lines. After all, in the UI, the accountant has to visually check everything and hit the OK button.

Therefore, I had to create a separate part to transform the OCR result into my internal structure:

  • block text;
  • score;
  • coordinates;
  • center by X/Y;
  • width and height.

After that, the blocks are grouped into lines: first sorting top to bottom, then merging vertically close elements, and within the line—sorting left to right.

Why is everything so complicated, you might ask? After all, a passport is a template!

Problem number two: thinking that text can only be parsed with regex

Yes, at first it seems: well, a passport is a template document. So, you can find dates, department codes, numbers, full names, and that's it.

But OCR tears logical blocks apart. For example, the issue date might be on the same line as the label "Issue Date," or it might be a line above or below.

The same goes for the date of birth, department code, gender, and place of birth. This is because when printing the passport, our template doesn't always lie perfectly flat.

Yes, there are boxes for each text. That's exactly what we need to find! So I had to create my own parser for each separate box and experiment with different arrangements of letters/numbers from various passports.

For example:

  • the issue date is searched next to the anchor "Issue Date";
  • the date of birth is searched by the anchor "born...", and if that doesn't work, the second found date is taken;
  • the department code is searched by the mask 000-000;
  • gender is searched by options MALE., MALE, FEMALE., FEMALE;
  • place of birth and "issued by" are collected separately because they can occupy multiple lines.

Problem number three: searching for full names not only by text but also by geometry

A separate pain point is the full name. I have no idea why it jumps in height across all passports and why it can't be printed strictly in millimeters on one line? Well, who am I to ask these questions?

So, for the full name, I had to add a search based on the geometry of the OCR boxes.

The logic is as follows:

  1. Find the OCR box with the field anchor: "LAST NAME," "FIRST NAME," "PATRONYMIC."
  2. Look at the candidates nearby.
  3. First, check the value to the right on the same line.
  4. Then check the value above the label.
  5. Discard junk, numbers, and service words.
  6. Choose the most similar option.

In fact, a passport is not just about text. It's about strict geometry. And essentially, we can use geometric search of OCR boxes everywhere.

Problem number four: series and number of the passport

Well, here everything is basically clear. The series and number are on the right side, and also vertically. That is, printed in a different orientation. Regular OCR tries to read it as horizontal text, and instead of numbers, you get some new letters of the alphabet.

So, for the series/number, I had to build a separate pipeline:

  1. Find the right vertical zone of the passport.
  2. Crop it as a separate image.
  3. Rotate the crop in both directions.
  4. Run OCR separately on these options.
  5. Try to assemble the series based on templates: two digits + two digits, or series next to the number.

The conclusion of this "tale" is that a passport cannot be properly recognized with a single OCR request. Different zones of the document behave differently, and some fields require separate processing.

Problem number five: the passport is there, the scanner is there, but no one said how to scan it correctly

You might think—it's a minor detail. Hysterical laughter.

It's easier for us to teach a machine to do a somersault with a passport than to give strict instructions to the accountant: scan strictly with the head up, one page of the passport on one page, preferably in good resolution.

So we write a preprocessor for each passport.

After the first recognition, the parser checks how many critical fields were not found. Critical fields are the issue date, date of birth, gender, full name, issued by, place of birth, etc. That is, all our main fields.

If there are too many such errors—more than half—then the document is likely:

  • upside down;
  • lying sideways;
  • very poorly readable;
  • or is it not a passport at all? O_o

Then we try to rotate the image by 180, 90, and even 270 degrees, recognize it again, and compare the results.

The best option is the one with fewer errors and more filled fields.

The final meaning is actually simple: if everything is completely bad on the first run, then we need to rotate the document.

Problem number six: slightly improving quality for the recognition engine

It's not exactly a problem, just a standard stage of the pipeline before shoving the photo into the OCR model.

For example: convert to grayscale, make it a bit more contrasty, increase sharpness, and some other filters that I tested and collected specifically for the passport.

The final preprocessor turned out to be:

  1. Conversion to grayscale.
  2. Resizing by the long side.
  3. Light noise reduction.
  4. Soft local contrast enhancement.
  5. Very careful unsharp.
  6. Saving in grayscale PNG.

So the goal was not to "make it pretty for humans," but to neatly improve readability for PaddleOCR and not break the letters.

The preprocessor is also not used all the time. First, the service tries regular recognition. If the document is flat, but there are still many critical errors, only then does fallback through preprocessing kick in.

Preprocessing is not always a blessing. Sometimes the original image is better understood by OCR than the "improved" one.

Logs are mandatory

And of course, logs are mandatory. OCR without logs is like reading tea leaves.

The accountant tells us that the passport was not recognized. But without logs, we won't understand what happened at all. So I had to attach JSON logs for key events:

  • module loading;
  • file normalization;
  • OCR launch;
  • line grouping;
  • field parsing;
  • rotation attempts;
  • preprocessor launch;
  • selection of the best result;
  • final summary.

This greatly simplifies debugging. Especially when you're not working with one perfect test passport, but with real photos that can be crooked, dark, with glare, etc.