Document OCR Platform & 1C Integration
Technology Stack
- Python
- PaddleOCR
- Flask
- Docker
- JSON
- 1C
- Nginx
Problem
Internal document workflows needed structured data rather than raw OCR text. The service had to recognize several internal document types, validate uncertain fields and pass reviewed results into 1C without relying on an external recognition service.
What I built
- Implemented PaddleOCR-based recognition for internal and foreign passports, SNILS, Russian tax IDs (INN) and patents.
- Added field normalization, validation, deduplication, error handling and operation logging.
- Built JSON and multipart interfaces with asynchronous delivery into 1C workflows.
Architecture
document image → image checks → PaddleOCR → field extraction and validation → review-ready structure → 1C integration
The Flask service was packaged with Docker and exposed internally through Nginx. Recognition and integration errors remained traceable.
Result
- Delivered an internal, locally operated OCR workflow for multiple document formats.
- Completed beta checks across the supported document variants and established a structured integration boundary for 1C.