Compare commits

..
3 Commits
14 changed files with 171 additions and 14 deletions
@@ -6,3 +6,4 @@ numpy
tenacity
pysubs2
pyannote.audio
python-dotenv
@@ -0,0 +1,33 @@
# Current State of AI Transcriber V2
**Date:** January 12, 2026
**Version:** 2.0 (Refactored & Robust)
## 🚀 Recently Completed Features
1. **Unified Translation Logic**: All scripts now use a shared fallback pipeline: **Primary (User Pref) -> Secondary -> Local LLM (Ollama) -> MyMemory**.
2. **Flattened Directory Structure**: Removed nested `ai_transcriber_v2/ai_transcriber_v2` folders. All V2 modules are now in the root of `ai_transcriber_v2/`.
3. **Flatpak & Bazzite Compatibility**:
* Added `flatpak-spawn --host` support for `ffmpeg`, `ffprobe`, and `ollama`.
* Updated `mount_truenas.sh` to handle UID/GID mapping for write permissions on remote shares.
4. **Hardware-Accelerated Safety Checks**:
* **Size Validation**: Rejects any remuxed file that drops more than 20% of the original size.
* **Duration Validation**: Rejects any remuxed file where the duration differs by more than 1 second.
* **Zero-Byte Check**: Deletes empty outputs immediately.
5. **Intelligent Skipping**:
* Whisper detects source language; if it matches the target (e.g., English to English), translation is skipped entirely.
6. **Progress Tracking**:
* Integrated `tqdm` progress bars for iterative translation steps.
* Enabled `verbose=True` for Whisper to show live transcription segments.
7. **Stability**:
* Fixed SQLAlchemy `DetachedInstanceError` by expunging objects from sessions in `tracker.py`.
* Added `GracefulKiller` for clean `Ctrl+C` shutdowns (finishes current file, then exits).
## 🛠 Active Setup
- **Environment**: Bazzite (Linux) running VS Code via Flatpak.
- **Local LLM**: Ollama with `llama3` (or `dolphin-llama3`).
- **Media Tools**: Host-level FFmpeg/FFprobe accessible via `flatpak-spawn`.
## 📌 Next Steps / Future Ideas
- Consider batching small SRT segments for Gemini to reduce API calls and improve context.
- Add a GUI or Web dashboard for tracking `job_history.db`.
- Implement auto-retry for specific "failed_validation" files with different model parameters.
@@ -0,0 +1,31 @@
# Technical Overview
## 🏗 Architecture
The project is modularized into specialized Python scripts:
* **`main.py`**: The entry point. Manages the batch processing loop and job tracking.
* **`extractor.py`**: Media handling via FFmpeg. Responsible for audio extraction and subtitle embedding. Contains the "Flatpak Escape" logic.
* **`transcriber.py`**: Integration with `openai-whisper`. Manages GPU health checks and model loading.
* **`translator.py`**: The AI translation engine. Implements the 4-tier fallback logic (Gemini -> Google -> Ollama -> MyMemory).
* **`utils.py`**: Shared utilities for encoding detection (chardet), port checking, and service health monitoring.
* **`tracker.py`**: Persistence layer using SQLite/SQLAlchemy to track job status across runs.
## 🛡 Safety Mechanisms
1. **Duration Match Check**: Uses `ffprobe` to ensure the final subbed video length matches the original source.
2. **File Size Sanity**: Rejects any remux operation that results in a file < 80% of the original size (preventing video stream loss).
3. **SRT Health Check**: Compares the last timestamp of the translated SRT against the original transcript to detect partial/truncated translations.
4. **Encoding Detection**: Uses `chardet` to reliably read foreign subtitle files without manual configuration.
## 🐳 Flatpak / Sandbox Support
Since this project is designed for Bazzite/Atomic distros, all system-level calls (`ffmpeg`, `ffprobe`, `ollama`) are wrapped in a check that detects the presence of `/.flatpak-info`. If found, it automatically prefixes commands with `flatpak-spawn --host` to utilize system-installed binaries.
## 💾 Database Schema
The `job_history.db` tracks:
- `file_path`: Absolute path to source.
- `status`: PENDING, PROCESSING, COMPLETED, FAILED.
- `step_status`: Individual status for Extract, Transcribe, Translate, and Embed steps.
- `error_message`: Captured stack traces for failed jobs.
@@ -0,0 +1,51 @@
# User Guide: AI Video Transcriber & Translator V2
This tool automates the process of extracting audio from videos, transcribing it using OpenAI Whisper, translating the text via AI (Gemini, Llama3, or Google), and embedding the results back into the video as soft subtitles.
## 📋 Prerequisites
1. **FFmpeg**: Must be installed on your host system.
* On Bazzite: `brew install ffmpeg`
2. **Ollama (Optional but Recommended)**: For private, local translation.
* Run `./ai_transcriber_v2/install_local_llm.sh`
* Pull a model: `ollama pull llama3`
3. **Python Packages**:
```bash
pip install -r ai_transcriber_v2/requirements.txt
```
## 🚀 How to Run
### 1. The Easy Way (Wizard)
Perfect for first-time runs or single folders.
```bash
./ai_transcriber_v2/run_wizard_v2.py
```
Follow the interactive prompts to set your languages, model size, and preferences.
### 2. The Power Way (CLI)
For advanced automation.
```bash
python3 ai_transcriber_v2/main.py /path/to/videos --lang English --prefer-local --embed --cleanup
```
### 3. The Library Fixer (Recovery)
If you have a folder with existing transcripts or partial translations that need fixing:
```bash
./ai_transcriber_v2/recover_and_fix_v2.py /path/to/folder
```
This script intelligently scans for missing translations or translations that don't match the video duration.
## 🛠 Features
* **Prefer Local LLM**: Use `--prefer-local` to prioritize your laptop's GPU (via Ollama) for all translations.
* **Safety First**: The script will NEVER replace your original video if the new one is significantly smaller or has a different duration.
* **Diarization**: Use `--diarize` to identify different speakers (requires HuggingFace token).
* **Graceful Exit**: Press `Ctrl+C` once to stop the script. It will finish the current file and save its progress before closing.
## 📂 File Naming Convention
- `video.srt`: Original language transcript.
- `video.English.srt`: Gemini translated subtitles.
- `video.English.deep_translate.srt`: Google Translate subtitles.
- `video.English.local_llm.srt`: Ollama translated subtitles.
- `video.subbed.mp4`: The final result with embedded soft-subs.
+16 -4
View File
@@ -1,6 +1,7 @@
import argparse
import os
import sys
import socket
from dotenv import load_dotenv
# Load environment variables from central .env_files directory
@@ -21,7 +22,7 @@ else:
from extractor import extract_audio, embed_subtitles
from transcriber import transcribe_audio, save_as_srt, load_whisper_model
from translator import translate_with_auto_fallback
from utils import validate_and_repair_srt, check_srt_duration_match, GracefulKiller, ensure_ollama_running, check_service_availability, check_path_permissions
from utils import validate_and_repair_srt, check_srt_duration_match, GracefulKiller, ensure_ollama_running, check_service_availability, check_path_permissions, LANGUAGE_MAP
from diarizer import diarize_audio, merge_diarization_with_transcript
import tracker
from tracker import JobStatus
@@ -75,6 +76,7 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
transcript_exists = os.path.exists(transcript_file) and not args.force
final_srt_path = transcript_file
detected_iso = None
if transcript_exists:
tracker.logger.info(f"Transcript exists: {transcript_file}. Skipping transcription.")
@@ -84,6 +86,7 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
# Use loaded_model if available
result = transcribe_audio(audio_path, model_size=args.model, language=source_lang, loaded_model=loaded_model)
segments = result["segments"]
detected_iso = result.get("language")
if args.diarize:
hf_token = args.hf_token or os.getenv("HF_TOKEN")
@@ -122,8 +125,16 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
translation_success = False
method_used = "None"
# Check existing
if (os.path.exists(base_translated) or os.path.exists(deep_translated) or os.path.exists(local_translated)) and not args.force:
# Check if source language matches target language
target_iso = LANGUAGE_MAP.get(args.lang)
if detected_iso and target_iso and detected_iso == target_iso:
tracker.logger.info(f"Source language '{detected_iso}' matches target '{target_iso}'. Skipping translation.")
final_srt_path = transcript_file
translation_success = True
method_used = "Source Match"
# Check existing (if not already handled by match)
elif (os.path.exists(base_translated) or os.path.exists(deep_translated) or os.path.exists(local_translated)) and not args.force:
if os.path.exists(local_translated):
translated_file = local_translated
method_used = "Local LLM (Existing)"
@@ -173,7 +184,8 @@ def process_file(file_path, args, source_lang=None, loaded_model=None, service_s
tracker.logger.error(f"VALIDATION FAILED: {msg}")
tracker.logger.error("Marking translation as failed due to incomplete coverage.")
redo_file = os.path.join(os.path.dirname(file_path), "redo_queue.txt")
hostname = socket.gethostname()
redo_file = os.path.join(os.path.dirname(file_path), f"redo_queue_{hostname}.txt")
with open(redo_file, "a", encoding="utf-8") as rf:
rf.write(f"{file_path} | {msg}\n")
@@ -3,6 +3,7 @@ import os
import sys
import argparse
import subprocess
import socket
from dotenv import load_dotenv
from datetime import datetime
import pysubs2
@@ -36,7 +37,8 @@ def process_recovery(folder_path, target_lang="English", prefer_deep=False, pref
else:
print("Preference: Gemini (API) > DeepTranslate")
recovery_log_file = os.path.join(folder_path, "recovery_status.log")
hostname = socket.gethostname()
recovery_log_file = os.path.join(folder_path, f"recovery_status_{hostname}.log")
print(f"Logging actions to: {recovery_log_file}")
# Ensure Ollama is ready
@@ -10,3 +10,4 @@ deep-translator
ollama
chardet
tqdm
python-dotenv
@@ -1,14 +1,18 @@
import logging
import os
import socket
from datetime import datetime
from sqlalchemy import create_engine, Column, Integer, String, DateTime, Enum, Text
from sqlalchemy.orm import declarative_base, sessionmaker
import enum
# Get Hostname for namespacing
HOSTNAME = socket.gethostname()
# Setup Logging
log_dir = "logs"
os.makedirs(log_dir, exist_ok=True)
log_file = os.path.join(log_dir, f"transcriber_{datetime.now().strftime('%Y%m%d')}.log")
log_file = os.path.join(log_dir, f"transcriber_{HOSTNAME}_{datetime.now().strftime('%Y%m%d')}.log")
logging.basicConfig(
level=logging.INFO,
@@ -22,7 +26,7 @@ logger = logging.getLogger(__name__)
# Database Setup
Base = declarative_base()
DB_FILE = "job_history.db"
DB_FILE = f"job_history_{HOSTNAME}.db"
class JobStatus(enum.Enum):
PENDING = "pending"
@@ -56,6 +60,10 @@ def get_job(file_path):
job = Job(file_path=file_path)
session.add(job)
session.commit()
session.refresh(job) # Ensure we have the ID and defaults
# Detach from session so we can use it after session.close()
session.expunge(job)
session.close()
return job
@@ -7,6 +7,7 @@ import pysubs2
from deep_translator import GoogleTranslator, MyMemoryTranslator
import ollama
from tqdm import tqdm
from utils import LANGUAGE_MAP
# Define a retry decorator
# ... (retry_policy remains)
@@ -21,6 +22,11 @@ def translate_via_ollama(source_srt_content, target_language="English", model="l
# Using tqdm for progress bar
for line in tqdm(subs, desc=" Ollama Progress", unit="line"):
text = line.text.strip()
# Skip empty, numeric-only, or extremely short non-word text
if not text or text.isdigit() or len(text) < 2:
continue
if text:
prompt = (
f"Translate this subtitle text to {target_language}. Output ONLY the translation.\n"
@@ -53,6 +59,11 @@ def translate_fallback_mymemory(source_srt_content, target_language="en"):
for line in tqdm(subs, desc=" MyMemory Progress", unit="line"):
text = line.text.strip()
# Skip empty, numeric-only, or extremely short non-word text
if not text or text.isdigit() or len(text) < 2:
continue
if text:
if len(text) > 500: # MyMemory has stricter limits often
continue
@@ -88,6 +99,11 @@ def translate_fallback_free(source_srt_content, target_language="en"):
# Simple line-by-line translation
for line in tqdm(subs, desc=" DeepTranslate Progress", unit="line"):
text = line.text.strip()
# Skip empty, numeric-only, or extremely short non-word text
if not text or text.isdigit() or len(text) < 2:
continue
if text:
# Sanity check: Skip lines that are too long
if len(text) > 4000:
@@ -243,13 +259,8 @@ def translate_with_auto_fallback(srt_content, target_language="English", prefer_
tuple: (translated_content, method_name) or (None, None) if all failed.
"""
# Map full language name to code for DeepTranslate
lang_map = {
"English": "en", "French": "fr", "Spanish": "es", "German": "de",
"Italian": "it", "Portuguese": "pt", "Russian": "ru",
"Japanese": "ja", "Chinese": "zh-CN"
}
target_code = lang_map.get(target_language, "en")
# Use central language mapping
target_code = LANGUAGE_MAP.get(target_language, "en")
# Determine which services to even try
def is_ok(name):
@@ -9,6 +9,13 @@ import shutil
import chardet
from tqdm import tqdm
# Global Language Mapping
LANGUAGE_MAP = {
"English": "en", "French": "fr", "Spanish": "es", "German": "de",
"Italian": "it", "Portuguese": "pt", "Russian": "ru",
"Japanese": "ja", "Chinese": "zh-CN", "auto": "auto"
}
class GracefulKiller:
"""
Handles SIGINT (Ctrl+C) and SIGTERM signals.