adding latest documentation and changes to the script.

This commit is contained in:
2026-01-12 11:47:39 -05:00
parent 7ad3715c52
commit 2a7878ddad
10 changed files with 144 additions and 10 deletions
@@ -0,0 +1,31 @@
# Technical Overview
## 🏗 Architecture
The project is modularized into specialized Python scripts:
* **`main.py`**: The entry point. Manages the batch processing loop and job tracking.
* **`extractor.py`**: Media handling via FFmpeg. Responsible for audio extraction and subtitle embedding. Contains the "Flatpak Escape" logic.
* **`transcriber.py`**: Integration with `openai-whisper`. Manages GPU health checks and model loading.
* **`translator.py`**: The AI translation engine. Implements the 4-tier fallback logic (Gemini -> Google -> Ollama -> MyMemory).
* **`utils.py`**: Shared utilities for encoding detection (chardet), port checking, and service health monitoring.
* **`tracker.py`**: Persistence layer using SQLite/SQLAlchemy to track job status across runs.
## 🛡 Safety Mechanisms
1. **Duration Match Check**: Uses `ffprobe` to ensure the final subbed video length matches the original source.
2. **File Size Sanity**: Rejects any remux operation that results in a file < 80% of the original size (preventing video stream loss).
3. **SRT Health Check**: Compares the last timestamp of the translated SRT against the original transcript to detect partial/truncated translations.
4. **Encoding Detection**: Uses `chardet` to reliably read foreign subtitle files without manual configuration.
## 🐳 Flatpak / Sandbox Support
Since this project is designed for Bazzite/Atomic distros, all system-level calls (`ffmpeg`, `ffprobe`, `ollama`) are wrapped in a check that detects the presence of `/.flatpak-info`. If found, it automatically prefixes commands with `flatpak-spawn --host` to utilize system-installed binaries.
## 💾 Database Schema
The `job_history.db` tracks:
- `file_path`: Absolute path to source.
- `status`: PENDING, PROCESSING, COMPLETED, FAILED.
- `step_status`: Individual status for Extract, Transcribe, Translate, and Embed steps.
- `error_message`: Captured stack traces for failed jobs.
@@ -0,0 +1,51 @@
# User Guide: AI Video Transcriber & Translator V2
This tool automates the process of extracting audio from videos, transcribing it using OpenAI Whisper, translating the text via AI (Gemini, Llama3, or Google), and embedding the results back into the video as soft subtitles.
## 📋 Prerequisites
1. **FFmpeg**: Must be installed on your host system.
* On Bazzite: `brew install ffmpeg`
2. **Ollama (Optional but Recommended)**: For private, local translation.
* Run `./ai_transcriber_v2/install_local_llm.sh`
* Pull a model: `ollama pull llama3`
3. **Python Packages**:
```bash
pip install -r ai_transcriber_v2/requirements.txt
```
## 🚀 How to Run
### 1. The Easy Way (Wizard)
Perfect for first-time runs or single folders.
```bash
./ai_transcriber_v2/run_wizard_v2.py
```
Follow the interactive prompts to set your languages, model size, and preferences.
### 2. The Power Way (CLI)
For advanced automation.
```bash
python3 ai_transcriber_v2/main.py /path/to/videos --lang English --prefer-local --embed --cleanup
```
### 3. The Library Fixer (Recovery)
If you have a folder with existing transcripts or partial translations that need fixing:
```bash
./ai_transcriber_v2/recover_and_fix_v2.py /path/to/folder
```
This script intelligently scans for missing translations or translations that don't match the video duration.
## 🛠 Features
* **Prefer Local LLM**: Use `--prefer-local` to prioritize your laptop's GPU (via Ollama) for all translations.
* **Safety First**: The script will NEVER replace your original video if the new one is significantly smaller or has a different duration.
* **Diarization**: Use `--diarize` to identify different speakers (requires HuggingFace token).
* **Graceful Exit**: Press `Ctrl+C` once to stop the script. It will finish the current file and save its progress before closing.
## 📂 File Naming Convention
- `video.srt`: Original language transcript.
- `video.English.srt`: Gemini translated subtitles.
- `video.English.deep_translate.srt`: Google Translate subtitles.
- `video.English.local_llm.srt`: Ollama translated subtitles.
- `video.subbed.mp4`: The final result with embedded soft-subs.