Engineering Work

AI Voice Assistant on NVIDIA Jetson

Technology Stack

  • Python
  • Flask
  • Faster Whisper
  • CUDA
  • NVIDIA Jetson
  • Docker
  • Local LLM
  • Nginx
  • Bitrix24 API

Problem

Online meetings needed a locally operated workflow for recording, transcription and follow-up AI processing. The solution had to run on available NVIDIA Jetson hardware and remain accessible as a secured internal service.

What I built

  • Created a Flask MVP that records meetings and transcribes audio with Faster Whisper.
  • Moved the pipeline to NVIDIA Jetson, enabled CUDA execution and packaged the service in Docker.
  • Added authenticated HTTPS access, TXT transcript export, local LLM processing and Bitrix24 integration.

Architecture

meeting audio → recording service → Faster Whisper transcription → stored transcript → local LLM / Bitrix24 workflow

Browsers expose microphone input only in a secure context, so the application was published over HTTPS. DNAT forwarded ports 80 and 443 to Nginx so an SSL certificate could be issued and the voice interface served over TLS. Containerization kept the speech and application dependencies reproducible on the edge device.

Result

  • Completed an end-to-end deployment and test on Jetson hardware.
  • Delivered a secured internal transcription service with downstream AI and business-system integration points.