Files
personal_development/video_transcription/ai_transcriber_v2/context/CURRENT_STATE.md
T

2.0 KiB

Current State of AI Transcriber V2

Date: January 12, 2026 Version: 2.0 (Refactored & Robust)

🚀 Recently Completed Features

  1. Unified Translation Logic: All scripts now use a shared fallback pipeline: Primary (User Pref) -> Secondary -> Local LLM (Ollama) -> MyMemory.
  2. Flattened Directory Structure: Removed nested ai_transcriber_v2/ai_transcriber_v2 folders. All V2 modules are now in the root of ai_transcriber_v2/.
  3. Flatpak & Bazzite Compatibility:
    • Added flatpak-spawn --host support for ffmpeg, ffprobe, and ollama.
    • Updated mount_truenas.sh to handle UID/GID mapping for write permissions on remote shares.
  4. Hardware-Accelerated Safety Checks:
    • Size Validation: Rejects any remuxed file that drops more than 20% of the original size.
    • Duration Validation: Rejects any remuxed file where the duration differs by more than 1 second.
    • Zero-Byte Check: Deletes empty outputs immediately.
  5. Intelligent Skipping:
    • Whisper detects source language; if it matches the target (e.g., English to English), translation is skipped entirely.
  6. Progress Tracking:
    • Integrated tqdm progress bars for iterative translation steps.
    • Enabled verbose=True for Whisper to show live transcription segments.
  7. Stability:
    • Fixed SQLAlchemy DetachedInstanceError by expunging objects from sessions in tracker.py.
    • Added GracefulKiller for clean Ctrl+C shutdowns (finishes current file, then exits).

🛠 Active Setup

  • Environment: Bazzite (Linux) running VS Code via Flatpak.
  • Local LLM: Ollama with llama3 (or dolphin-llama3).
  • Media Tools: Host-level FFmpeg/FFprobe accessible via flatpak-spawn.

📌 Next Steps / Future Ideas

  • Consider batching small SRT segments for Gemini to reduce API calls and improve context.
  • Add a GUI or Web dashboard for tracking job_history.db.
  • Implement auto-retry for specific "failed_validation" files with different model parameters.