2.0 KiB
2.0 KiB
Current State of AI Transcriber V2
Date: January 12, 2026 Version: 2.0 (Refactored & Robust)
🚀 Recently Completed Features
- Unified Translation Logic: All scripts now use a shared fallback pipeline: Primary (User Pref) -> Secondary -> Local LLM (Ollama) -> MyMemory.
- Flattened Directory Structure: Removed nested
ai_transcriber_v2/ai_transcriber_v2folders. All V2 modules are now in the root ofai_transcriber_v2/. - Flatpak & Bazzite Compatibility:
- Added
flatpak-spawn --hostsupport forffmpeg,ffprobe, andollama. - Updated
mount_truenas.shto handle UID/GID mapping for write permissions on remote shares.
- Added
- Hardware-Accelerated Safety Checks:
- Size Validation: Rejects any remuxed file that drops more than 20% of the original size.
- Duration Validation: Rejects any remuxed file where the duration differs by more than 1 second.
- Zero-Byte Check: Deletes empty outputs immediately.
- Intelligent Skipping:
- Whisper detects source language; if it matches the target (e.g., English to English), translation is skipped entirely.
- Progress Tracking:
- Integrated
tqdmprogress bars for iterative translation steps. - Enabled
verbose=Truefor Whisper to show live transcription segments.
- Integrated
- Stability:
- Fixed SQLAlchemy
DetachedInstanceErrorby expunging objects from sessions intracker.py. - Added
GracefulKillerfor cleanCtrl+Cshutdowns (finishes current file, then exits).
- Fixed SQLAlchemy
🛠 Active Setup
- Environment: Bazzite (Linux) running VS Code via Flatpak.
- Local LLM: Ollama with
llama3(ordolphin-llama3). - Media Tools: Host-level FFmpeg/FFprobe accessible via
flatpak-spawn.
📌 Next Steps / Future Ideas
- Consider batching small SRT segments for Gemini to reduce API calls and improve context.
- Add a GUI or Web dashboard for tracking
job_history.db. - Implement auto-retry for specific "failed_validation" files with different model parameters.