13 KiB
VideoDetect - Video Classification System
A production-grade video classification system for processing large-scale video corpora (~30TB) to classify videos based on demographic presence using deep learning models.
Architecture
┌─────────────┐ ┌──────────────┐ ┌─────────────┐
│ Scanner │────▶│ Processor │────▶│ Results │
│ (STORY-02) │ │ (STORY-03/04)│ │ (STORY-05) │
└─────────────┘ └──────────────┘ └─────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌──────────────┐ ┌─────────────┐
│ MariaDB │◀───▶│ GPU (P40) │◀───▶│ Export │
│ (Metadata) │ │ (Inference) │ │ (Parquet) │
└─────────────┘ └──────────────┘ └─────────────┘
│
▼
┌─────────────┐ ┌──────────────┐
│ Review UI │◀───▶│ Active Learn │
│ (STORY-06) │ │ (STORY-07) │
└─────────────┘ └──────────────┘
│
▼
┌─────────────┐
│ Monitoring │
│ (STORY-08) │
└─────────────┘
Quick Start
Prerequisites
- Docker & Docker Compose v2+
- NVIDIA Container Toolkit
- 2× Tesla P40 GPUs (or compatible)
- NAS/SMB mount at
/mnt/nas(configurable via env vars)
Environment Variables
# Database
export DB_ROOT_PASSWORD=your_root_password
export DB_PASSWORD=your_db_password
# Paths (optional, defaults shown)
export NAS_INPUT_PATH=/mnt/nas/input
export NAS_OUTPUT_PATH=/mnt/nas/output
export MODELS_PATH=/mnt/nas/models
export TRAINING_PATH=/mnt/nas/training
# Grafana admin password
export GRAFANA_PASSWORD=your_grafana_password
Start the Stack
docker-compose up -d
Verify Services
# Check all containers are running
docker-compose ps
# Check GPU visibility in worker
docker exec videodetect-worker nvidia-smi
# Check database connectivity
docker exec videodetect-mariadb mysql -u videodetect -p videodetect -e "SHOW TABLES;"
# Access UI
open http://localhost:5000
# Access Grafana
open http://localhost:3000 (admin / your_grafana_password)
# Access Prometheus
open http://localhost:9090
System Operation
How Processes Start
Service Initialization:
- MariaDB starts first with health check
- Worker initializes via
src/main.py:- Loads
config.yaml - Sets up JSON logging with rotation
- Connects to MariaDB (connection pooling)
- Initializes database schema
- Verifies GPU availability (CUDA/PyTorch)
- Starts the
DirectoryScannerin a background thread (scans permanent storage in place) - Creates
WorkerPoolwith 1 worker thread - Enters job processing loop
- Loads
- UI starts Flask review interface via Gunicorn (2 workers)
- Monitoring starts Prometheus and Grafana independently
Processing Pipeline
The worker follows this flow for each video:
Pending → Lock → Probe → Sample → Detect → Classify → Aggregate → Route → Persist → Export → Cleanup → Completed
Detailed Steps:
-
Job Queue - Atomically lock
PENDINGjobs viaUPDATE status = 'PROCESSING'- Priority: newest files first (
last_scan_time DESC) - Max concurrent: 1 per GPU
- Priority: newest files first (
-
Probe Video - Extract metadata via FFprobe
- Duration, codec, resolution
- Validate against codec whitelist (H.264, H.265, VP8/9, AV1)
- Mark
UNSCANNABLEif invalid
-
Sample Frames - Extract frames at configured interval (default: 30s)
- Save as JPEG to
/scratch/{video_id}/frames/ - Quality: 2 (lower=better)
- Save as JPEG to
-
Detect Faces - YOLOv8n TensorRT inference (FP32)
- Batch size auto-tuned by GPU memory monitor
- NMS filtering (IoU: 0.45, confidence: 0.25)
- Cap: 10 faces/frame, 100 faces/video
-
Extract Crops - Resize detected faces to 224×224
- Save to
/scratch/{video_id}/crops/
- Save to
-
Classify Crops - MobileNetV3-Small TensorRT inference
- Temperature-scaled softmax (T=1.0)
- Returns confidence per crop
-
Aggregate Confidence - Combine crop confidences into video-level score
- Strategy:
max(most conservative) - Alternatives:
weighted_mean,top_k_mean
- Strategy:
-
Route Decision - Threshold-based routing:
C ≥ 0.75→ MATCH0.45 ≤ C < 0.75→ REVIEW (human annotation)C < 0.45→ SKIP- No faces → SKIP
-
Persist Results - Atomic transaction:
- Update
videostable (confidence, routing, status) - Insert
processing_logsrow (audit trail) - State guard: only update if
status='PROCESSING'
- Update
-
Export - Buffer and batch export (default: 100 videos)
- Format: Parquet with Snappy compression
- Path:
/data/output/{model_version}/ - Fallback: JSONL if Parquet fails
-
Cleanup - Delete
/scratch/{video_id}/directory- Only after successful persistence
- Prevents orphaned scratch files
Directory Scanner Service (runs alongside the worker):
- Scans the permanent storage location in place (no staging/copy step)
- Walks
/data/inputevery 2 hours by default (configurable viascanner.scan_interval_seconds) - Detects new, modified, and removed video files by comparing against the DB
- Filters to video files by extension
- Computes SHA256 hash, probes metadata, validates codec
- Queues any video that has not been scanned yet as
PENDINGfor the worker pool - Single-instance guard: an in-process lock plus a DB lock (with a lease) ensure only one scan runs at a time — a long-running scan never overlaps another, even across multiple worker replicas. The lock lease is refreshed via heartbeats during the scan and is taken over automatically if a scanner crashes.
Configuration Reference
All configuration is in config.yaml. Environment variable override format: VD_<SECTION>_<KEY> (e.g., VD_SAMPLING_INTERVAL_SECONDS=60).
Key Configuration Sections
Sampling & Thresholds:
sampling:
interval_seconds: 30 # Frame extraction frequency
quality: 2 # JPEG quality (1-31, lower=better)
format: jpeg
thresholds:
T_high: 0.75 # MATCH threshold
T_low: 0.45 # REVIEW threshold
GPU & Batching:
gpu:
max_memory_gb: 18 # Target VRAM usage
batch_size: auto # Auto-tune based on available VRAM
batching:
max_batch_size: 16 # Maximum batch size
vram_target_gb: 16 # Target VRAM for batch tuning
vram_reduce_threshold_gb: 16 # Reduce batch if above
vram_increase_threshold_gb: 10 # Increase batch if below
Storage Paths:
storage:
scratch_path: /scratch # Temporary processing (tmpfs)
input_path: /data/input # Source videos (NFS)
output_path: /data/output # Results (local/NAS)
models_path: /models # TensorRT models
training_path: /data/training # Training data
Database:
database:
host: mariadb
port: 3306
name: videodetect
user: videodetect
password: videodetect123
pool_size: 20 # Connection pool size
pool_min: 5
pool_recycle: 3600 # Recycle connections after 1h
Face Detection:
face_detection:
model: yolo8n
model_path: /models/face_detector/face_detector.trt
input_size: 640
confidence_threshold: 0.25
iou_threshold: 0.45
max_faces_per_frame: 10
max_faces_per_video: 100
Classification & Aggregation:
classifier:
model: mobilenetv3-small
model_path: /models/classifier/classifier.trt
input_size: 224
temperature: 1.0 # Calibration temperature
aggregation:
strategy: max # max, weighted_mean, top_k_mean
alpha: 1.0 # weighted_mean weight for mean
beta: 0.1 # weighted_mean weight for variance
top_k: 3 # top_k_mean: average top 3 scores
Export:
export:
format: parquet # parquet, jsonl, or both
compression: snappy
batch_size: 100 # Export after N videos
include_frame_confidences: true
Review UI:
review_ui:
host: "0.0.0.0"
port: 5000
per_page: 20 # Pagination
top_k_frames: 5 # Show top-k contributing frames
auth_enabled: false # No auth per TC-06
ssl_enabled: false # Internal LAN only
Volume Mounts
From docker-compose.yml:
- Input: NFS mount →
/data/input(read-only) - Output:
./output→/data/output - Models:
./models→/models - Training:
./training→/data/training - Scratch: 100GB tmpfs at
/scratch(RAM disk)
Key Design Principles
- Atomic state transitions - Database locks prevent race conditions
- Crash recovery -
PROCESSINGjobs automatically requeued on restart - Idempotent - Re-running same video produces same result
- Stateless - Scratch cleanup after each job
- Fail-safe - 3 retry attempts before marking
ERROR - No auth/SSL - Internal LAN deployment per TC-06
Project Structure
VideoDetect/
├── docker-compose.yml # Multi-service orchestration
├── config.yaml # All configuration
├── db/
│ └── schema.sql # MariaDB schema
├── worker/
│ ├── Dockerfile # Worker container image
│ └── requirements.txt # Python dependencies
├── ui/
│ ├── Dockerfile # UI container image
│ └── review/ # Review UI templates
├── src/
│ ├── main.py # Worker entry point
│ ├── db_connector.py # Database connection layer
│ ├── config_loader.py # Configuration management
│ ├── logging_config.py # Logging setup
│ ├── scanner.py # Directory scanner (STORY-02)
│ ├── prober.py # Video probing (STORY-02)
│ ├── frame_sampler.py # Frame extraction (STORY-03)
│ ├── face_detector.py # Face detection (STORY-04)
│ ├── classifier.py # Classification (STORY-05)
│ ├── aggregator.py # Confidence aggregation (STORY-05)
│ ├── router.py # Threshold routing (STORY-05)
│ ├── result_updater.py # Result persistence (STORY-06)
│ ├── data_export.py # Parquet/JSONL export (STORY-06)
│ ├── review_api.py # Review API (STORY-07)
│ ├── active_learning/ # Active learning pipeline (STORY-08)
│ │ ├── label_ingestor.py
│ │ ├── trainer.py
│ │ ├── validator.py
│ │ └── registry.py
│ ├── metrics.py # Prometheus metrics (STORY-09)
│ ├── crash_recovery.py # Crash recovery (STORY-09)
│ └── drift_detector.py # Drift detection (STORY-09)
├── monitoring/
│ ├── prometheus/
│ │ └── prometheus.yml # Prometheus config
│ └── grafana/
│ ├── dashboards/ # Grafana dashboard JSONs
│ └── provisioning/ # Grafana data source config
├── STORY-01.md through STORY-09.md # Implementation stories
├── Plan.md # Implementation plan
└── Requirements.md # Requirements document
Implementation Stories
| Story | Description | Sprint | Points |
|---|---|---|---|
| STORY-01 | Foundation & Infrastructure | 1-2 | 13 |
| STORY-02 | Core Ingestion & Codec Handling | 3-4 | 21 |
| STORY-03 | Frame Sampling | 5 | 13 |
| STORY-04 | Face Detection | 5-6 | 21 |
| STORY-05 | Classification & Confidence | 5-6 | 21 |
| STORY-06 | Results Persistence & Export | 5-6 | 13 |
| STORY-07 | Review Interface | 7 | 21 |
| STORY-08 | Active Learning Pipeline | 7-8 | 34 |
| STORY-09 | Observability & Hardening | 9+ | 34 |
Hardware Requirements
- GPUs: 2× Tesla P40 24GB (compute capability 5.2)
- CUDA: ≤ 11.8
- PyTorch: ≤ 2.1.0
- Inference: FP32 only (no Tensor Cores)
- Storage: NVMe/SSD for scratch, NAS/SMB for input/output
- RAM: ≥ 32GB system RAM
Key Design Decisions
- TensorRT FP32 only — Tesla P40 has no Tensor Cores; FP16 would not provide benefit
- YOLOv8n + MobileNetV3 — Lightweight models optimized for throughput over accuracy
- Max aggregation — Most conservative confidence strategy; use highest frame confidence
- Temperature scaling — Calibrates confidence scores without retraining
- Head-only fine-tuning — Faster retraining with frozen backbone
- Parquet export — Efficient columnar format for analytics
- No auth/SSL — Per TC-06, internal LAN only
License
Proprietary — VideoDetect Project