Files
viral_project/openshorts/CLAUDE.md
T

4.5 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

OpenShorts is an AI-powered vertical video generator that transforms long YouTube videos or local uploads into viral-ready short clips (9:16 format) for TikTok, Instagram Reels, and YouTube Shorts. Uses Google Gemini 2.0 Flash for viral moment detection and title generation.

Development Commands

Local Development (Docker)

docker compose up --build   # Build and run full stack

Frontend Only (Dashboard)

cd dashboard
npm install
npm run dev       # Dev server with HMR (port 5173)
npm run build     # Production build
npm run lint      # ESLint (strict, --max-warnings 0)

Backend Only

pip install -r requirements.txt
uvicorn app:app --host 0.0.0.0 --port 8000

Architecture

Core Processing Pipeline

  1. Ingest - YouTube download (yt-dlp) or local upload
  2. Transcription - faster-whisper with word-level timestamps
  3. Scene Detection - PySceneDetect for segment boundaries
  4. AI Analysis - Gemini identifies 3-15 viral moments (15-60 sec each)
  5. FFmpeg Extraction - Precise clip cutting
  6. AI Cropping - Vertical reframing with subject tracking
  7. Effects/Subtitles - Optional AI-generated FFmpeg filters
  8. Hook Overlay - Text overlays with styled fonts
  9. Voice Dubbing - Optional ElevenLabs AI translation (30+ languages)
  10. S3 Backup - Silent background upload
  11. Social Distribution - Upload-Post API (async upload)

Key Files

File Purpose
main.py Core video processing: transcription, scene detection, clip extraction, vertical reframing
app.py FastAPI server with async job queue and REST endpoints
editor.py Gemini AI integration for dynamic video effects (FFmpeg filter generation)
hooks.py Hook text overlay generation with font rendering
s3_uploader.py AWS S3 upload with caching
subtitles.py SRT generation, FFmpeg subtitle burning, and dubbed video transcription
translate.py ElevenLabs dubbing API for AI voice translation
dashboard/src/App.jsx Main React component with state management
dashboard/src/components/TranslateModal.jsx Voice dubbing UI with language selection

Dual-Mode Video Reframing

  • TRACK Mode (single subject): MediaPipe face detection + YOLOv8 fallback with "Heavy Tripod" stabilization
  • GENERAL Mode (groups/landscapes): Blurred background layout preserving full width

Key Classes

  • SmoothedCameraman - Stabilized camera movement with safe zone logic (prevents jitter)
  • SpeakerTracker - Prevents rapid speaker switching, handles temporary occlusions

API Endpoints

Method Route Purpose
POST /api/process Submit video for processing
GET /api/status/{job_id} Poll job status and logs
POST /api/edit Apply AI video effects
POST /api/subtitle Generate and apply subtitles (auto-transcribes dubbed videos)
POST /api/hook Add text hook overlays
POST /api/translate AI voice dubbing via ElevenLabs
GET /api/translate/languages List supported dubbing languages
POST /api/social/post Post to social media (async upload)

Concurrency Model

Async job queue with semaphore-based concurrency control. Configure via MAX_CONCURRENT_JOBS env var (default: 5). Jobs auto-cleanup after 1 hour.

Environment Variables

Server-side (.env):

  • AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION, AWS_S3_BUCKET - For S3 backup
  • MAX_CONCURRENT_JOBS - Concurrent processing limit (default: 5)
  • VITE_API_URL - Production API URL override

Client-side (localStorage, encrypted):

  • GEMINI_API_KEY - Google Gemini API key (required)
  • ELEVENLABS_API_KEY - ElevenLabs API key for voice dubbing (optional)
  • UPLOAD_POST_API_KEY - Upload-Post API key for social posting (optional)

API keys are stored encrypted in the browser and sent via headers only when needed. Never stored server-side.

Tech Stack

  • Backend: Python 3.11, FastAPI, google-genai, faster-whisper, ultralytics (YOLOv8), mediapipe, opencv-python, yt-dlp, FFmpeg, httpx
  • Frontend: React 18, Vite 4, Tailwind CSS 3.4
  • External APIs: Google Gemini, ElevenLabs Dubbing, Upload-Post
  • Infrastructure: Docker + Docker Compose, AWS S3