4.5 KiB
4.5 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
OpenShorts is an AI-powered vertical video generator that transforms long YouTube videos or local uploads into viral-ready short clips (9:16 format) for TikTok, Instagram Reels, and YouTube Shorts. Uses Google Gemini 2.0 Flash for viral moment detection and title generation.
Development Commands
Local Development (Docker)
docker compose up --build # Build and run full stack
- Backend: http://localhost:8000 (FastAPI/Uvicorn)
- Frontend: http://localhost:5175 (Vite proxies API calls to backend)
Frontend Only (Dashboard)
cd dashboard
npm install
npm run dev # Dev server with HMR (port 5173)
npm run build # Production build
npm run lint # ESLint (strict, --max-warnings 0)
Backend Only
pip install -r requirements.txt
uvicorn app:app --host 0.0.0.0 --port 8000
Architecture
Core Processing Pipeline
- Ingest - YouTube download (yt-dlp) or local upload
- Transcription - faster-whisper with word-level timestamps
- Scene Detection - PySceneDetect for segment boundaries
- AI Analysis - Gemini identifies 3-15 viral moments (15-60 sec each)
- FFmpeg Extraction - Precise clip cutting
- AI Cropping - Vertical reframing with subject tracking
- Effects/Subtitles - Optional AI-generated FFmpeg filters
- Hook Overlay - Text overlays with styled fonts
- Voice Dubbing - Optional ElevenLabs AI translation (30+ languages)
- S3 Backup - Silent background upload
- Social Distribution - Upload-Post API (async upload)
Key Files
| File | Purpose |
|---|---|
main.py |
Core video processing: transcription, scene detection, clip extraction, vertical reframing |
app.py |
FastAPI server with async job queue and REST endpoints |
editor.py |
Gemini AI integration for dynamic video effects (FFmpeg filter generation) |
hooks.py |
Hook text overlay generation with font rendering |
s3_uploader.py |
AWS S3 upload with caching |
subtitles.py |
SRT generation, FFmpeg subtitle burning, and dubbed video transcription |
translate.py |
ElevenLabs dubbing API for AI voice translation |
dashboard/src/App.jsx |
Main React component with state management |
dashboard/src/components/TranslateModal.jsx |
Voice dubbing UI with language selection |
Dual-Mode Video Reframing
- TRACK Mode (single subject): MediaPipe face detection + YOLOv8 fallback with "Heavy Tripod" stabilization
- GENERAL Mode (groups/landscapes): Blurred background layout preserving full width
Key Classes
SmoothedCameraman- Stabilized camera movement with safe zone logic (prevents jitter)SpeakerTracker- Prevents rapid speaker switching, handles temporary occlusions
API Endpoints
| Method | Route | Purpose |
|---|---|---|
| POST | /api/process |
Submit video for processing |
| GET | /api/status/{job_id} |
Poll job status and logs |
| POST | /api/edit |
Apply AI video effects |
| POST | /api/subtitle |
Generate and apply subtitles (auto-transcribes dubbed videos) |
| POST | /api/hook |
Add text hook overlays |
| POST | /api/translate |
AI voice dubbing via ElevenLabs |
| GET | /api/translate/languages |
List supported dubbing languages |
| POST | /api/social/post |
Post to social media (async upload) |
Concurrency Model
Async job queue with semaphore-based concurrency control. Configure via MAX_CONCURRENT_JOBS env var (default: 5). Jobs auto-cleanup after 1 hour.
Environment Variables
Server-side (.env):
AWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEY,AWS_REGION,AWS_S3_BUCKET- For S3 backupMAX_CONCURRENT_JOBS- Concurrent processing limit (default: 5)VITE_API_URL- Production API URL override
Client-side (localStorage, encrypted):
GEMINI_API_KEY- Google Gemini API key (required)ELEVENLABS_API_KEY- ElevenLabs API key for voice dubbing (optional)UPLOAD_POST_API_KEY- Upload-Post API key for social posting (optional)
API keys are stored encrypted in the browser and sent via headers only when needed. Never stored server-side.
Tech Stack
- Backend: Python 3.11, FastAPI, google-genai, faster-whisper, ultralytics (YOLOv8), mediapipe, opencv-python, yt-dlp, FFmpeg, httpx
- Frontend: React 18, Vite 4, Tailwind CSS 3.4
- External APIs: Google Gemini, ElevenLabs Dubbing, Upload-Post
- Infrastructure: Docker + Docker Compose, AWS S3