9.8 KiB
Implementation Plan: Video Classification System
Overview
This document outlines the phased implementation plan for the Video Classification System. The plan breaks down the requirements into logical work packages (Sprints/Phases) that can be implemented, tested, and integrated sequentially.
Key Dependencies:
- Infrastructure must be available before any processing logic can run.
- Database Schema must be finalized early to support metadata logging.
- Model Training Data is a bottleneck; the initial model must be ready or a placeholder used for pipeline testing.
Phase 1: Foundation & Infrastructure (Sprint 1-2)
Goal: Establish the Docker environment, Database schema, and basic connectivity. No video processing yet, just the "skeleton."
1.1 Environment Setup
- Docker Compose Structure: Create
docker-compose.ymldefining services for:mariadb: Persistent volume for metadata.worker: Build context for PyTorch/TensorRT environment. Includenvidia-container-toolkitconfig for GPU passthrough.ui: Placeholder container (e.g., nginx serving a static "Under Construction" page) to verify network connectivity.
- Base Images:
- Create a custom Dockerfile for the worker based on
nvidia/cuda:11.8.0-runtime-ubuntu20.04. - Install PyTorch (1.13+ with CUDA 11.8), OpenCV, FFmpeg-python, and TensorRT prerequisites.
- Create a custom Dockerfile for the worker based on
- Volume Mounts:
- Map local NVMe path to
/scratchin worker container (tmpfs preferred for speed). - Map NAS/SMB path to
/data/inputand/data/outputin worker container. - Map
/modelsand/data/trainingfor persistence.
- Map local NVMe path to
1.2 Database Schema Design
- MariaDB Schema: Define tables in SQL migration script:
videos:id,file_path,file_hash,resolution_w,resolution_h,codec,status(NEW, PENDING, PROCESSING, COMPLETED, UNSCANNABLE, ERROR),last_scan_time,last_processed_time.processing_logs:id,video_id,model_version,frame_count,confidence_score,routing_decision,processed_at.models:version,status(ACTIVE, CANDIDATE, ARCHIVED),path,calibration_temp.
- Connection Layer: Implement a Python module (
db_connector.py) with connection pooling to MariaDB.
1.3 Configuration Management
- Config File: Create
config.yamlfor thresholds (T_high,T_low), sampling intervals, and GPU settings. - Logging Setup: Configure structured logging (JSON format) to stdout for Docker capture.
Exit Criteria:
docker-compose upstarts DB and Worker containers successfully.- Worker can connect to DB and create tables.
- GPU is visible inside the Worker container (
nvidia-smi).
Phase 2: Core Ingestion & Codec Handling (Sprint 3-4)
Goal: Implement the directory scanner, video probing, and robust error handling for codecs/resolutions.
2.1 Directory Scanner
- Scanner Service: Implement a background thread or separate script that walks
/data/input. - File Detection: Identify new files (not in DB) and update status to
PENDING. - Deduplication: Use
file_hash(SHA-256 of first 1MB or full file if small) to prevent re-processing identical files.
2.2 Video Probing & Metadata Extraction
- FFmpeg Probe: For each
PENDINGfile, runffprobeto extract:- Codec name.
- Resolution (width/height).
- Duration.
- Codec Validation: Maintain a whitelist of supported codecs (H.264, H.265, VP8, VP9, etc.).
- If codec is unsupported or
ffprobefails: Set status toUNSCANNABLE. - Log the specific error code/reason.
- If codec is unsupported or
- Resolution Handling: Store resolution in DB. Implement logic to downscale high-res frames (4K) if necessary to fit VRAM constraints during inference.
2.3 Batch Orchestration Skeleton
- Job Queue: Implement a simple priority queue or DB-based locking mechanism to assign files to Worker processes.
- State Management: Ensure atomic transitions from
PENDING→PROCESSINGto avoid duplicate processing.
Exit Criteria:
- New files in
/data/inputappear in DB with correct metadata within 60 seconds. - Unsupported codec files are marked
UNSCANNABLEwithout crashing the scanner. - DB accurately reflects the current state of the 30TB corpus.
Phase 3: Inference Pipeline & Model Integration (Sprint 5-6)
Goal: Implement frame sampling, face detection, classification, and confidence aggregation.
3.1 Frame Sampling
- FFmpeg Extraction: Implement logic to extract 1 frame per 30 seconds uniformly.
- Handle variable FPS gracefully.
- Save frames to
/scratch/tmp/with naming convention{video_id}_{timestamp}.jpg.
- Memory Management: Ensure frames are deleted from scratch space immediately after processing the video to prevent filling NVMe.
3.2 Face Detection
- Model Integration: Load YOLOv8n (or chosen lightweight detector) via ONNX Runtime or TensorRT.
- Batching: Implement dynamic batching for face detection to maximize GPU utilization.
- Crop Generation: Extract face crops, resize to inference input size (e.g., 224x224).
3.3 Classification & Aggregation
- Model Integration: Load MobileNetV3 (or chosen classifier) via TensorRT FP32.
- Inference: Run classification on face crops.
- Aggregation Logic:
- Calculate frame-level confidence
p_i. - Apply video-level aggregation (e.g.,
max(p_i)or weighted mean). - Apply temperature scaling for calibration.
- Calculate frame-level confidence
- Routing Decision: Compare final confidence
CagainstT_highandT_low. AssignMATCH,REVIEW, orSKIP.
3.4 Result Persistence
- DB Update: Update
videostable withprocessed_at,model_version, and statusCOMPLETED. - Log Entry: Insert record into
processing_logs. - Export: Write summary JSON/Parquet file to
/data/outputfor analytics.
Exit Criteria:
- A sample video is processed end-to-end: Frames extracted → Faces detected → Classified → Confidence calculated → DB updated.
- GPU memory usage stays under 18GB per GPU.
- Throughput meets baseline (>30 videos/hour/GPU on test set).
Phase 4: Active Learning & Review UI (Sprint 7-8)
Goal: Enable human-in-the-loop correction and model retraining.
4.1 Review UI
- UI Setup: Deploy Label Studio or custom Flask UI.
- Data Feed: Query DB for videos with routing_decision =
REVIEW. - Frontend: Display video player + top-k contributing frames. Allow annotator to toggle label (True/False) or correct classification.
- Export: Generate CSV/JSON export of annotated data with ground truth.
4.2 Fine-Tuning Pipeline
- Data Loader: Script to ingest exported annotations and prepare dataset for training.
- Training Job:
- Head-only fine-tuning on face crops.
- Use AdamW, LR=1e-3, Early Stopping.
- Save checkpoint to
/models/candidate/.
- Validation:
- Run candidate model on held-out validation set.
- Calculate F1 score and ECE (Expected Calibration Error).
- If ΔF1 > 0.02 and ECE < 0.08, promote model to
ACTIVE.
4.3 Model Swapping
- Registry Update: Update DB
modelstable to mark new version asACTIVEand old asARCHIVED. - Hot Reload: Restart Worker containers or signal process to reload new TensorRT engine.
Exit Criteria:
- Annotator can label a review video.
- Labeled data triggers a fine-tuning job.
- New model is validated and deployed automatically.
Phase 5: Observability, Monitoring & Hardening (Sprint 9+)
Goal: Ensure system reliability, debuggability, and long-term stability.
5.1 Metrics & Logging
- Prometheus Integration: Expose metrics:
videos_processed_total(by routing decision).gpu_utilization,gpu_memory_used.queue_depth(PENDING/PROCESSING counts).confidence_distributionhistogram.
- Grafana Dashboard: Create dashboards for:
- Throughput (videos/hour).
- Error rates (UNSCANNABLE/ERROR counts).
- Confidence drift alerts.
5.2 Fault Tolerance & Resumption
- Crash Recovery: Implement checkpointing. If Worker dies, mark
PROCESSINGjobs asPENDINGagain. - Idempotency: Ensure re-processing a file does not duplicate DB entries or outputs.
- Retry Logic: Auto-retry transient errors (e.g., network blip during NAS access) up to 3 times.
5.3 Drift Detection
- Weekly Job: Compare current confidence distribution to baseline.
- Alerting: Send alert if
p(C > 0.5)shifts >10% or Review Queue grows unbounded.
Exit Criteria:
- Dashboard shows real-time health of the system.
- System recovers gracefully from forced Worker termination.
- Alerts trigger correctly on simulated drift/errors.
Risk Mitigation
| Risk | Mitigation Strategy |
|---|---|
| GPU OOM | Implement strict batch size caps; downscale 4K frames to 1080p/720p before inference. |
| Slow NAS I/O | Cache frames in local NVMe tmpfs; minimize disk writes until final export. |
| Codec Variance | Extensive unit testing of ffprobe logic against known "bad" files. Graceful degradation to UNSCANNABLE. |
| Model Regression | Strict validation gate before model swap. Keep previous model in ARCHIVED state for quick rollback. |
| 30TB Scan Time | Incremental scanning using file_hash and last_modified timestamps. Avoid full rescan. |
Deliverables Checklist
docker-compose.yml- Worker Dockerfile & Requirements.txt
- MariaDB SQL Schema
- Python Source Code (
src/)scanner.pyprocessor.pymodel_runner.pydb_handler.pytrain.py
- Config Files (
config.yaml) - Grafana Dashboard JSON
- Documentation (
README.md,API_DOCS.md)