Files
personal_development/video_transcription/ai_transcriber_v2/documentation/TECHNICAL_OVERVIEW.md
T

1.9 KiB

Technical Overview

🏗 Architecture

The project is modularized into specialized Python scripts:

  • main.py: The entry point. Manages the batch processing loop and job tracking.
  • extractor.py: Media handling via FFmpeg. Responsible for audio extraction and subtitle embedding. Contains the "Flatpak Escape" logic.
  • transcriber.py: Integration with openai-whisper. Manages GPU health checks and model loading.
  • translator.py: The AI translation engine. Implements the 4-tier fallback logic (Gemini -> Google -> Ollama -> MyMemory).
  • utils.py: Shared utilities for encoding detection (chardet), port checking, and service health monitoring.
  • tracker.py: Persistence layer using SQLite/SQLAlchemy to track job status across runs.

🛡 Safety Mechanisms

  1. Duration Match Check: Uses ffprobe to ensure the final subbed video length matches the original source.
  2. File Size Sanity: Rejects any remux operation that results in a file < 80% of the original size (preventing video stream loss).
  3. SRT Health Check: Compares the last timestamp of the translated SRT against the original transcript to detect partial/truncated translations.
  4. Encoding Detection: Uses chardet to reliably read foreign subtitle files without manual configuration.

🐳 Flatpak / Sandbox Support

Since this project is designed for Bazzite/Atomic distros, all system-level calls (ffmpeg, ffprobe, ollama) are wrapped in a check that detects the presence of /.flatpak-info. If found, it automatically prefixes commands with flatpak-spawn --host to utilize system-installed binaries.

💾 Database Schema

The job_history.db tracks:

  • file_path: Absolute path to source.
  • status: PENDING, PROCESSING, COMPLETED, FAILED.
  • step_status: Individual status for Extract, Transcribe, Translate, and Embed steps.
  • error_message: Captured stack traces for failed jobs.