Engineering Work

Document OCR Platform & 1C Integration

Technology Stack

  • Python
  • PaddleOCR
  • Flask
  • Docker
  • JSON
  • 1C
  • Nginx

Problem

Internal document workflows needed structured data rather than raw OCR text. The service had to recognize several internal document types, validate uncertain fields and pass reviewed results into 1C without relying on an external recognition service.

What I built

  • Implemented PaddleOCR-based recognition for internal and foreign passports, SNILS, Russian tax IDs (INN) and patents.
  • Added field normalization, validation, deduplication, error handling and operation logging.
  • Built JSON and multipart interfaces with asynchronous delivery into 1C workflows.

Architecture

document image → image checks → PaddleOCR → field extraction and validation → review-ready structure → 1C integration

The Flask service was packaged with Docker and exposed internally through Nginx. Recognition and integration errors remained traceable.

Result

  • Delivered an internal, locally operated OCR workflow for multiple document formats.
  • Completed beta checks across the supported document variants and established a structured integration boundary for 1C.