adding latest documentation and changes to the script.

This commit is contained in:
2026-01-12 11:47:39 -05:00
parent 7ad3715c52
commit 2a7878ddad
10 changed files with 144 additions and 10 deletions
@@ -6,3 +6,4 @@ numpy
tenacity tenacity
pysubs2 pysubs2
pyannote.audio pyannote.audio
python-dotenv
@@ -0,0 +1,33 @@
# Current State of AI Transcriber V2
**Date:** January 12, 2026
**Version:** 2.0 (Refactored & Robust)
## 🚀 Recently Completed Features
1. **Unified Translation Logic**: All scripts now use a shared fallback pipeline: **Primary (User Pref) -> Secondary -> Local LLM (Ollama) -> MyMemory**.
2. **Flattened Directory Structure**: Removed nested `ai_transcriber_v2/ai_transcriber_v2` folders. All V2 modules are now in the root of `ai_transcriber_v2/`.
3. **Flatpak & Bazzite Compatibility**:
* Added `flatpak-spawn --host` support for `ffmpeg`, `ffprobe`, and `ollama`.
* Updated `mount_truenas.sh` to handle UID/GID mapping for write permissions on remote shares.
4. **Hardware-Accelerated Safety Checks**:
* **Size Validation**: Rejects any remuxed file that drops more than 20% of the original size.
* **Duration Validation**: Rejects any remuxed file where the duration differs by more than 1 second.
* **Zero-Byte Check**: Deletes empty outputs immediately.
5. **Intelligent Skipping**:
* Whisper detects source language; if it matches the target (e.g., English to English), translation is skipped entirely.
6. **Progress Tracking**:
* Integrated `tqdm` progress bars for iterative translation steps.
* Enabled `verbose=True` for Whisper to show live transcription segments.
7. **Stability**:
* Fixed SQLAlchemy `DetachedInstanceError` by expunging objects from sessions in `tracker.py`.
* Added `GracefulKiller` for clean `Ctrl+C` shutdowns (finishes current file, then exits).
## 🛠 Active Setup
- **Environment**: Bazzite (Linux) running VS Code via Flatpak.
- **Local LLM**: Ollama with `llama3` (or `dolphin-llama3`).
- **Media Tools**: Host-level FFmpeg/FFprobe accessible via `flatpak-spawn`.
## 📌 Next Steps / Future Ideas
- Consider batching small SRT segments for Gemini to reduce API calls and improve context.
- Add a GUI or Web dashboard for tracking `job_history.db`.
- Implement auto-retry for specific "failed_validation" files with different model parameters.
@@ -0,0 +1,31 @@
# Technical Overview
## 🏗 Architecture
The project is modularized into specialized Python scripts:
* **`main.py`**: The entry point. Manages the batch processing loop and job tracking.
* **`extractor.py`**: Media handling via FFmpeg. Responsible for audio extraction and subtitle embedding. Contains the "Flatpak Escape" logic.
* **`transcriber.py`**: Integration with `openai-whisper`. Manages GPU health checks and model loading.
* **`translator.py`**: The AI translation engine. Implements the 4-tier fallback logic (Gemini -> Google -> Ollama -> MyMemory).
* **`utils.py`**: Shared utilities for encoding detection (chardet), port checking, and service health monitoring.
* **`tracker.py`**: Persistence layer using SQLite/SQLAlchemy to track job status across runs.
## 🛡 Safety Mechanisms
1. **Duration Match Check**: Uses `ffprobe` to ensure the final subbed video length matches the original source.
2. **File Size Sanity**: Rejects any remux operation that results in a file < 80% of the original size (preventing video stream loss).
3. **SRT Health Check**: Compares the last timestamp of the translated SRT against the original transcript to detect partial/truncated translations.
4. **Encoding Detection**: Uses `chardet` to reliably read foreign subtitle files without manual configuration.
## 🐳 Flatpak / Sandbox Support
Since this project is designed for Bazzite/Atomic distros, all system-level calls (`ffmpeg`, `ffprobe`, `ollama`) are wrapped in a check that detects the presence of `/.flatpak-info`. If found, it automatically prefixes commands with `flatpak-spawn --host` to utilize system-installed binaries.
## 💾 Database Schema
The `job_history.db` tracks:
- `file_path`: Absolute path to source.
- `status`: PENDING, PROCESSING, COMPLETED, FAILED.
- `step_status`: Individual status for Extract, Transcribe, Translate, and Embed steps.
- `error_message`: Captured stack traces for failed jobs.
@@ -0,0 +1,51 @@
# User Guide: AI Video Transcriber & Translator V2
This tool automates the process of extracting audio from videos, transcribing it using OpenAI Whisper, translating the text via AI (Gemini, Llama3, or Google), and embedding the results back into the video as soft subtitles.
## 📋 Prerequisites
1. **FFmpeg**: Must be installed on your host system.
* On Bazzite: `brew install ffmpeg`
2. **Ollama (Optional but Recommended)**: For private, local translation.
* Run `./ai_transcriber_v2/install_local_llm.sh`
* Pull a model: `ollama pull llama3`
3. **Python Packages**:
```bash
pip install -r ai_transcriber_v2/requirements.txt
```
## 🚀 How to Run
### 1. The Easy Way (Wizard)
Perfect for first-time runs or single folders.
```bash
./ai_transcriber_v2/run_wizard_v2.py
```
Follow the interactive prompts to set your languages, model size, and preferences.
### 2. The Power Way (CLI)
For advanced automation.
```bash
python3 ai_transcriber_v2/main.py /path/to/videos --lang English --prefer-local --embed --cleanup
```
### 3. The Library Fixer (Recovery)
If you have a folder with existing transcripts or partial translations that need fixing:
```bash
./ai_transcriber_v2/recover_and_fix_v2.py /path/to/folder
```
This script intelligently scans for missing translations or translations that don't match the video duration.
## 🛠 Features
* **Prefer Local LLM**: Use `--prefer-local` to prioritize your laptop's GPU (via Ollama) for all translations.
* **Safety First**: The script will NEVER replace your original video if the new one is significantly smaller or has a different duration.
* **Diarization**: Use `--diarize` to identify different speakers (requires HuggingFace token).
* **Graceful Exit**: Press `Ctrl+C` once to stop the script. It will finish the current file and save its progress before closing.
## 📂 File Naming Convention
- `video.srt`: Original language transcript.
- `video.English.srt`: Gemini translated subtitles.
- `video.English.deep_translate.srt`: Google Translate subtitles.
- `video.English.local_llm.srt`: Ollama translated subtitles.
- `video.subbed.mp4`: The final result with embedded soft-subs.
+13 -3
View File
@@ -21,7 +21,7 @@ else:
from extractor import extract_audio, embed_subtitles from extractor import extract_audio, embed_subtitles
from transcriber import transcribe_audio, save_as_srt, load_whisper_model from transcriber import transcribe_audio, save_as_srt, load_whisper_model
from translator import translate_with_auto_fallback from translator import translate_with_auto_fallback
from utils import validate_and_repair_srt, check_srt_duration_match, GracefulKiller, ensure_ollama_running, check_service_availability, check_path_permissions from utils import validate_and_repair_srt, check_srt_duration_match, GracefulKiller, ensure_ollama_running, check_service_availability, check_path_permissions, LANGUAGE_MAP
from diarizer import diarize_audio, merge_diarization_with_transcript from diarizer import diarize_audio, merge_diarization_with_transcript
import tracker import tracker
from tracker import JobStatus from tracker import JobStatus
@@ -75,6 +75,7 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
transcript_exists = os.path.exists(transcript_file) and not args.force transcript_exists = os.path.exists(transcript_file) and not args.force
final_srt_path = transcript_file final_srt_path = transcript_file
detected_iso = None
if transcript_exists: if transcript_exists:
tracker.logger.info(f"Transcript exists: {transcript_file}. Skipping transcription.") tracker.logger.info(f"Transcript exists: {transcript_file}. Skipping transcription.")
@@ -84,6 +85,7 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
# Use loaded_model if available # Use loaded_model if available
result = transcribe_audio(audio_path, model_size=args.model, language=source_lang, loaded_model=loaded_model) result = transcribe_audio(audio_path, model_size=args.model, language=source_lang, loaded_model=loaded_model)
segments = result["segments"] segments = result["segments"]
detected_iso = result.get("language")
if args.diarize: if args.diarize:
hf_token = args.hf_token or os.getenv("HF_TOKEN") hf_token = args.hf_token or os.getenv("HF_TOKEN")
@@ -122,8 +124,16 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
translation_success = False translation_success = False
method_used = "None" method_used = "None"
# Check existing # Check if source language matches target language
if (os.path.exists(base_translated) or os.path.exists(deep_translated) or os.path.exists(local_translated)) and not args.force: target_iso = LANGUAGE_MAP.get(args.lang)
if detected_iso and target_iso and detected_iso == target_iso:
tracker.logger.info(f"Source language '{detected_iso}' matches target '{target_iso}'. Skipping translation.")
final_srt_path = transcript_file
translation_success = True
method_used = "Source Match"
# Check existing (if not already handled by match)
elif (os.path.exists(base_translated) or os.path.exists(deep_translated) or os.path.exists(local_translated)) and not args.force:
if os.path.exists(local_translated): if os.path.exists(local_translated):
translated_file = local_translated translated_file = local_translated
method_used = "Local LLM (Existing)" method_used = "Local LLM (Existing)"
@@ -10,3 +10,4 @@ deep-translator
ollama ollama
chardet chardet
tqdm tqdm
python-dotenv
@@ -56,6 +56,10 @@ def get_job(file_path):
job = Job(file_path=file_path) job = Job(file_path=file_path)
session.add(job) session.add(job)
session.commit() session.commit()
session.refresh(job) # Ensure we have the ID and defaults
# Detach from session so we can use it after session.close()
session.expunge(job)
session.close() session.close()
return job return job
@@ -7,6 +7,7 @@ import pysubs2
from deep_translator import GoogleTranslator, MyMemoryTranslator from deep_translator import GoogleTranslator, MyMemoryTranslator
import ollama import ollama
from tqdm import tqdm from tqdm import tqdm
from utils import LANGUAGE_MAP
# Define a retry decorator # Define a retry decorator
# ... (retry_policy remains) # ... (retry_policy remains)
@@ -243,13 +244,8 @@ def translate_with_auto_fallback(srt_content, target_language="English", prefer_
tuple: (translated_content, method_name) or (None, None) if all failed. tuple: (translated_content, method_name) or (None, None) if all failed.
""" """
# Map full language name to code for DeepTranslate # Use central language mapping
lang_map = { target_code = LANGUAGE_MAP.get(target_language, "en")
"English": "en", "French": "fr", "Spanish": "es", "German": "de",
"Italian": "it", "Portuguese": "pt", "Russian": "ru",
"Japanese": "ja", "Chinese": "zh-CN"
}
target_code = lang_map.get(target_language, "en")
# Determine which services to even try # Determine which services to even try
def is_ok(name): def is_ok(name):
@@ -9,6 +9,13 @@ import shutil
import chardet import chardet
from tqdm import tqdm from tqdm import tqdm
# Global Language Mapping
LANGUAGE_MAP = {
"English": "en", "French": "fr", "Spanish": "es", "German": "de",
"Italian": "it", "Portuguese": "pt", "Russian": "ru",
"Japanese": "ja", "Chinese": "zh-CN", "auto": "auto"
}
class GracefulKiller: class GracefulKiller:
""" """
Handles SIGINT (Ctrl+C) and SIGTERM signals. Handles SIGINT (Ctrl+C) and SIGTERM signals.