adding latest documentation and changes to the script.
This commit is contained in:
@@ -6,3 +6,4 @@ numpy
|
||||
tenacity
|
||||
pysubs2
|
||||
pyannote.audio
|
||||
python-dotenv
|
||||
Binary file not shown.
@@ -0,0 +1,33 @@
|
||||
# Current State of AI Transcriber V2
|
||||
|
||||
**Date:** January 12, 2026
|
||||
**Version:** 2.0 (Refactored & Robust)
|
||||
|
||||
## 🚀 Recently Completed Features
|
||||
1. **Unified Translation Logic**: All scripts now use a shared fallback pipeline: **Primary (User Pref) -> Secondary -> Local LLM (Ollama) -> MyMemory**.
|
||||
2. **Flattened Directory Structure**: Removed nested `ai_transcriber_v2/ai_transcriber_v2` folders. All V2 modules are now in the root of `ai_transcriber_v2/`.
|
||||
3. **Flatpak & Bazzite Compatibility**:
|
||||
* Added `flatpak-spawn --host` support for `ffmpeg`, `ffprobe`, and `ollama`.
|
||||
* Updated `mount_truenas.sh` to handle UID/GID mapping for write permissions on remote shares.
|
||||
4. **Hardware-Accelerated Safety Checks**:
|
||||
* **Size Validation**: Rejects any remuxed file that drops more than 20% of the original size.
|
||||
* **Duration Validation**: Rejects any remuxed file where the duration differs by more than 1 second.
|
||||
* **Zero-Byte Check**: Deletes empty outputs immediately.
|
||||
5. **Intelligent Skipping**:
|
||||
* Whisper detects source language; if it matches the target (e.g., English to English), translation is skipped entirely.
|
||||
6. **Progress Tracking**:
|
||||
* Integrated `tqdm` progress bars for iterative translation steps.
|
||||
* Enabled `verbose=True` for Whisper to show live transcription segments.
|
||||
7. **Stability**:
|
||||
* Fixed SQLAlchemy `DetachedInstanceError` by expunging objects from sessions in `tracker.py`.
|
||||
* Added `GracefulKiller` for clean `Ctrl+C` shutdowns (finishes current file, then exits).
|
||||
|
||||
## 🛠 Active Setup
|
||||
- **Environment**: Bazzite (Linux) running VS Code via Flatpak.
|
||||
- **Local LLM**: Ollama with `llama3` (or `dolphin-llama3`).
|
||||
- **Media Tools**: Host-level FFmpeg/FFprobe accessible via `flatpak-spawn`.
|
||||
|
||||
## 📌 Next Steps / Future Ideas
|
||||
- Consider batching small SRT segments for Gemini to reduce API calls and improve context.
|
||||
- Add a GUI or Web dashboard for tracking `job_history.db`.
|
||||
- Implement auto-retry for specific "failed_validation" files with different model parameters.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Technical Overview
|
||||
|
||||
## 🏗 Architecture
|
||||
|
||||
The project is modularized into specialized Python scripts:
|
||||
|
||||
* **`main.py`**: The entry point. Manages the batch processing loop and job tracking.
|
||||
* **`extractor.py`**: Media handling via FFmpeg. Responsible for audio extraction and subtitle embedding. Contains the "Flatpak Escape" logic.
|
||||
* **`transcriber.py`**: Integration with `openai-whisper`. Manages GPU health checks and model loading.
|
||||
* **`translator.py`**: The AI translation engine. Implements the 4-tier fallback logic (Gemini -> Google -> Ollama -> MyMemory).
|
||||
* **`utils.py`**: Shared utilities for encoding detection (chardet), port checking, and service health monitoring.
|
||||
* **`tracker.py`**: Persistence layer using SQLite/SQLAlchemy to track job status across runs.
|
||||
|
||||
## 🛡 Safety Mechanisms
|
||||
|
||||
1. **Duration Match Check**: Uses `ffprobe` to ensure the final subbed video length matches the original source.
|
||||
2. **File Size Sanity**: Rejects any remux operation that results in a file < 80% of the original size (preventing video stream loss).
|
||||
3. **SRT Health Check**: Compares the last timestamp of the translated SRT against the original transcript to detect partial/truncated translations.
|
||||
4. **Encoding Detection**: Uses `chardet` to reliably read foreign subtitle files without manual configuration.
|
||||
|
||||
## 🐳 Flatpak / Sandbox Support
|
||||
|
||||
Since this project is designed for Bazzite/Atomic distros, all system-level calls (`ffmpeg`, `ffprobe`, `ollama`) are wrapped in a check that detects the presence of `/.flatpak-info`. If found, it automatically prefixes commands with `flatpak-spawn --host` to utilize system-installed binaries.
|
||||
|
||||
## 💾 Database Schema
|
||||
|
||||
The `job_history.db` tracks:
|
||||
- `file_path`: Absolute path to source.
|
||||
- `status`: PENDING, PROCESSING, COMPLETED, FAILED.
|
||||
- `step_status`: Individual status for Extract, Transcribe, Translate, and Embed steps.
|
||||
- `error_message`: Captured stack traces for failed jobs.
|
||||
@@ -0,0 +1,51 @@
|
||||
# User Guide: AI Video Transcriber & Translator V2
|
||||
|
||||
This tool automates the process of extracting audio from videos, transcribing it using OpenAI Whisper, translating the text via AI (Gemini, Llama3, or Google), and embedding the results back into the video as soft subtitles.
|
||||
|
||||
## 📋 Prerequisites
|
||||
|
||||
1. **FFmpeg**: Must be installed on your host system.
|
||||
* On Bazzite: `brew install ffmpeg`
|
||||
2. **Ollama (Optional but Recommended)**: For private, local translation.
|
||||
* Run `./ai_transcriber_v2/install_local_llm.sh`
|
||||
* Pull a model: `ollama pull llama3`
|
||||
3. **Python Packages**:
|
||||
```bash
|
||||
pip install -r ai_transcriber_v2/requirements.txt
|
||||
```
|
||||
|
||||
## 🚀 How to Run
|
||||
|
||||
### 1. The Easy Way (Wizard)
|
||||
Perfect for first-time runs or single folders.
|
||||
```bash
|
||||
./ai_transcriber_v2/run_wizard_v2.py
|
||||
```
|
||||
Follow the interactive prompts to set your languages, model size, and preferences.
|
||||
|
||||
### 2. The Power Way (CLI)
|
||||
For advanced automation.
|
||||
```bash
|
||||
python3 ai_transcriber_v2/main.py /path/to/videos --lang English --prefer-local --embed --cleanup
|
||||
```
|
||||
|
||||
### 3. The Library Fixer (Recovery)
|
||||
If you have a folder with existing transcripts or partial translations that need fixing:
|
||||
```bash
|
||||
./ai_transcriber_v2/recover_and_fix_v2.py /path/to/folder
|
||||
```
|
||||
This script intelligently scans for missing translations or translations that don't match the video duration.
|
||||
|
||||
## 🛠 Features
|
||||
|
||||
* **Prefer Local LLM**: Use `--prefer-local` to prioritize your laptop's GPU (via Ollama) for all translations.
|
||||
* **Safety First**: The script will NEVER replace your original video if the new one is significantly smaller or has a different duration.
|
||||
* **Diarization**: Use `--diarize` to identify different speakers (requires HuggingFace token).
|
||||
* **Graceful Exit**: Press `Ctrl+C` once to stop the script. It will finish the current file and save its progress before closing.
|
||||
|
||||
## 📂 File Naming Convention
|
||||
- `video.srt`: Original language transcript.
|
||||
- `video.English.srt`: Gemini translated subtitles.
|
||||
- `video.English.deep_translate.srt`: Google Translate subtitles.
|
||||
- `video.English.local_llm.srt`: Ollama translated subtitles.
|
||||
- `video.subbed.mp4`: The final result with embedded soft-subs.
|
||||
@@ -21,7 +21,7 @@ else:
|
||||
from extractor import extract_audio, embed_subtitles
|
||||
from transcriber import transcribe_audio, save_as_srt, load_whisper_model
|
||||
from translator import translate_with_auto_fallback
|
||||
from utils import validate_and_repair_srt, check_srt_duration_match, GracefulKiller, ensure_ollama_running, check_service_availability, check_path_permissions
|
||||
from utils import validate_and_repair_srt, check_srt_duration_match, GracefulKiller, ensure_ollama_running, check_service_availability, check_path_permissions, LANGUAGE_MAP
|
||||
from diarizer import diarize_audio, merge_diarization_with_transcript
|
||||
import tracker
|
||||
from tracker import JobStatus
|
||||
@@ -75,6 +75,7 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
|
||||
transcript_exists = os.path.exists(transcript_file) and not args.force
|
||||
|
||||
final_srt_path = transcript_file
|
||||
detected_iso = None
|
||||
|
||||
if transcript_exists:
|
||||
tracker.logger.info(f"Transcript exists: {transcript_file}. Skipping transcription.")
|
||||
@@ -84,6 +85,7 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
|
||||
# Use loaded_model if available
|
||||
result = transcribe_audio(audio_path, model_size=args.model, language=source_lang, loaded_model=loaded_model)
|
||||
segments = result["segments"]
|
||||
detected_iso = result.get("language")
|
||||
|
||||
if args.diarize:
|
||||
hf_token = args.hf_token or os.getenv("HF_TOKEN")
|
||||
@@ -122,8 +124,16 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
|
||||
translation_success = False
|
||||
method_used = "None"
|
||||
|
||||
# Check existing
|
||||
if (os.path.exists(base_translated) or os.path.exists(deep_translated) or os.path.exists(local_translated)) and not args.force:
|
||||
# Check if source language matches target language
|
||||
target_iso = LANGUAGE_MAP.get(args.lang)
|
||||
if detected_iso and target_iso and detected_iso == target_iso:
|
||||
tracker.logger.info(f"Source language '{detected_iso}' matches target '{target_iso}'. Skipping translation.")
|
||||
final_srt_path = transcript_file
|
||||
translation_success = True
|
||||
method_used = "Source Match"
|
||||
|
||||
# Check existing (if not already handled by match)
|
||||
elif (os.path.exists(base_translated) or os.path.exists(deep_translated) or os.path.exists(local_translated)) and not args.force:
|
||||
if os.path.exists(local_translated):
|
||||
translated_file = local_translated
|
||||
method_used = "Local LLM (Existing)"
|
||||
|
||||
@@ -10,3 +10,4 @@ deep-translator
|
||||
ollama
|
||||
chardet
|
||||
tqdm
|
||||
python-dotenv
|
||||
@@ -56,6 +56,10 @@ def get_job(file_path):
|
||||
job = Job(file_path=file_path)
|
||||
session.add(job)
|
||||
session.commit()
|
||||
session.refresh(job) # Ensure we have the ID and defaults
|
||||
|
||||
# Detach from session so we can use it after session.close()
|
||||
session.expunge(job)
|
||||
session.close()
|
||||
return job
|
||||
|
||||
|
||||
@@ -7,6 +7,7 @@ import pysubs2
|
||||
from deep_translator import GoogleTranslator, MyMemoryTranslator
|
||||
import ollama
|
||||
from tqdm import tqdm
|
||||
from utils import LANGUAGE_MAP
|
||||
|
||||
# Define a retry decorator
|
||||
# ... (retry_policy remains)
|
||||
@@ -243,13 +244,8 @@ def translate_with_auto_fallback(srt_content, target_language="English", prefer_
|
||||
tuple: (translated_content, method_name) or (None, None) if all failed.
|
||||
"""
|
||||
|
||||
# Map full language name to code for DeepTranslate
|
||||
lang_map = {
|
||||
"English": "en", "French": "fr", "Spanish": "es", "German": "de",
|
||||
"Italian": "it", "Portuguese": "pt", "Russian": "ru",
|
||||
"Japanese": "ja", "Chinese": "zh-CN"
|
||||
}
|
||||
target_code = lang_map.get(target_language, "en")
|
||||
# Use central language mapping
|
||||
target_code = LANGUAGE_MAP.get(target_language, "en")
|
||||
|
||||
# Determine which services to even try
|
||||
def is_ok(name):
|
||||
|
||||
@@ -9,6 +9,13 @@ import shutil
|
||||
import chardet
|
||||
from tqdm import tqdm
|
||||
|
||||
# Global Language Mapping
|
||||
LANGUAGE_MAP = {
|
||||
"English": "en", "French": "fr", "Spanish": "es", "German": "de",
|
||||
"Italian": "it", "Portuguese": "pt", "Russian": "ru",
|
||||
"Japanese": "ja", "Chinese": "zh-CN", "auto": "auto"
|
||||
}
|
||||
|
||||
class GracefulKiller:
|
||||
"""
|
||||
Handles SIGINT (Ctrl+C) and SIGTERM signals.
|
||||
|
||||
Reference in New Issue
Block a user