Zadachkin — OCR + LLM Problem-Solving Platform
Technology Stack
- Python
- PaddleOCR
- Flask
- Telegram API
- Ollama
- LLM APIs
- Docker
Product
Zadachkin processes photographed tasks and turns them into structured text that an LLM can interpret and solve. The main challenge was creating a dependable boundary between noisy user images, OCR and model output.
What I built
- Developed a Flask OCR API around PaddleOCR with MIME, image-size and timeout validation.
- Added hash-based caching, request logging and a single endpoint for recognized text.
- Built a Telegram bot that accepts photos or files and sends recognized content to an LLM service.
- Tested both API models and a local Ollama/Qwen path with filtered Markdown or LaTeX output.
Architecture
Telegram image → validation → PaddleOCR service → normalized text → local or API LLM → formatted answer
The OCR component and bot were containerized independently, keeping recognition replaceable and simplifying local experiments.
Result
- Delivered an end-to-end OCR-to-LLM workflow for photographed tasks.