# Video Classification System: Requirements Document
**Version:** 1.2
**Author:** AI Architecture Consultant
**Target Audience:** Engineering Leads, ML Engineers, DevOps, Product/Project Managers
**Purpose:** Define functional, technical, and operational requirements to drive implementation story creation and
sprint planning.
---
## 1. System Overview & Objectives
The system processes a large corpus of video files (~30TB) as a local background batch job to classify each video into two
categories:
- `MATCH`: Video contains at least one Black male subject
- `NO_MATCH`: Video does not contain the target demographic
The system prioritizes **throughput and speed** over maximum accuracy, outputs **calibrated confidence scores**,
routes low-confidence results to a **manual review queue**, and supports an **active learning loop** to
incrementally improve model performance using reviewed samples.
---
## 2. Functional Requirements
| ID | Requirement | Priority | Notes |
|----|-------------|----------|-------|
| FR-01 | Configurable frame sampling interval (default: 1 frame per 30 seconds) | P0 | Must support override per
job/batch. Uniform temporal sampling preferred. |
| FR-02 | Face detection on all sampled frames | P0 | Lightweight detector only; no full-body or scene analysis. |
| FR-03 | Binary demographic classification of detected face crops | P0 | Outputs probability `p ∈ [0,1]` for
target class. |
| FR-04 | Video-level confidence aggregation & threshold routing | P0 | Aggregates frame-level scores → video
confidence `C`. Routes to `MATCH`, `REVIEW`, or `SKIP`. |
| FR-05 | Manual review interface for low-confidence videos | P1 | Displays video + contributing frames/crops +
model confidence. Supports binary labeling. |
| FR-06 | Active learning pipeline (label ingestion → fine-tuning → deployment) | P1 | Batch retraining only. No
online learning. Versioned model swaps. |
| FR-07 | Metadata logging & audit trail | P1 | Stores video ID, timestamps, frame counts, confidence scores,
routing decision, model version. |
| FR-08 | Batch job orchestration & crash recovery | P0 | Supports resume, parallel GPU scheduling, and
deterministic IDempotent processing. |
| FR-09 | Directory Scanning & Sync | P0 | Process to scan input directories, detect new/removed files, and sync state to MariaDB. |
| FR-10 | Codec & Resolution Detection & Handling | P0 | Detect video properties; handle unsupported codecs by flagging files as `UNSCANNABLE`. |
---
## 3. Non-Functional Requirements
| ID | Requirement | Target | Notes |
|----|-------------|--------|-------|
| NFR-01 | Throughput | ≥ 30 videos/hour/GPU (≈ 60 videos/hour total) | Baseline; tunable via sampling interval &
batch size. |
| NFR-02 | Latency per video | ≤ 45 seconds end-to-end (15-min avg video) | Excludes I/O bottlenecks; measured at
compute stage. |
| NFR-03 | GPU Memory Safety | ≤ 18GB per GPU sustained | Leaves headroom for OS, queues, and peak allocation. |
| NFR-04 | Determinism & Reproducibility | Config-seeded randomness, versioned models | Enables auditability and
rollback. |
| NFR-05 | Fault Tolerance | Auto-retry on transient failures; skip & log on fatal errors | Prevents batch
poisoning. |
| NFR-06 | Observability | Prometheus/Grafana metrics + structured logging | Tracks FPS, queue depth, confidence
distribution, drift alerts. |
| NFR-07 | Data Volume Handling | Efficient indexing for ~30TB dataset | Metadata stored in MariaDB; file existence verified via hashing/checksums if needed. |
| NFR-08 | Codec Agnosticism | Handle H.264, H.265, VP8, VP9, AV1, MJPEG, etc. | Unsupported codecs flagged gracefully. |
---
## 4. Technical & Environmental Constraints
| ID | Constraint | Details |
|----|------------|---------|
| TC-01 | Hardware | 2× Tesla P40 24GB (compute capability 5.2, PCIe 3.0, **no Tensor Cores**) |
| TC-02 | CUDA/Torch Compatibility | CUDA ≤ 11.8, PyTorch ≤ 2.1.0, FP32 inference only |
| TC-03 | Storage I/O | Fast local NVMe/SSD for temp frame cache; shared NAS/SMB for video input/output. 30TB Capacity. |
| TC-04 | Framework Stack | PyTorch → ONNX → TensorRT FP32; FFmpeg/OpenCV for sampling; **MariaDB** for metadata/state |
| TC-05 | Deployment Model | **Docker Compose** orchestrates all services (Workers, DB, UI). GPUs passed via `nvidia-container-toolkit`. |
| TC-06 | Network Security | Internal LAN only. No reverse proxy, SSL, or auth required for UI. |
---
## 5. Data & Storage Architecture
| Layer | Specification |
|-------|---------------|
| **Input** | Raw video files (MP4, MOV, AVI, MKV) totaling ~30TB. Growth is slow (archival nature). Resolutions: 480p–4K. Codecs: Mixed. |
| **Scratch** | `tmpfs` or fast local SSD for extracted frames & face crops. Auto-cleaned post-job. |
| **Metadata Store** | **MariaDB** database storing: `file_path`, `file_hash`, `resolution`, `codec`, `last_scan_time`, `last_processed_time`, `model_version_used`, `confidence_score`, `routing_decision`, `status` (OK, UNSCANNABLE, ERROR). |
| **Results** | Parquet/JSON lines exported for analytics: `{video_id, model_version, sample_count, confidence_scores[], video_confidence, routing, processed_at}` |
| **Training Data** | Versioned directory structure: `/data/v1/crops/`, `/data/v1/labels/`. Augmentation pipeline
applied at training time. |
| **Model Registry** | `/models/` with semantic versioning. Active, candidate, and archived states tracked. |
---
## 6. Confidence Scoring & Routing Logic
| Stage | Specification |
|-------|---------------|
| **Frame-Level** | Raw logits → temperature-scaled softmax → calibrated `p_i ∈ [0,1]` |
| **Video-Level** | `C = max(p_i)` OR `C = softmax(α·mean(p_i) + β·var(p_i))` (configurable) |
| **Routing Thresholds** | `C ≥ T_high` → `MATCH`
`T_low ≤ C < T_high` → `REVIEW`
`C < T_low` → `SKIP`
*(Default: T_high=0.75, T_low=0.45)* |
| **Calibration** | Temperature scaling evaluated on held-out set every training cycle. Stores `T` with model
checkpoint. |
| **Error Handling** | If FFmpeg/OpenCV fails to decode frame or detect codec: Set status to `UNSCANNABLE`. Log error. Do not retry indefinitely. |
---
## 7. Active Learning Pipeline
| Component | Specification |
|-----------|---------------|
| **Review Queue** | Lightweight web UI (Label Studio or custom Flask/FastAPI). Shows video player, top-k
contributing frames, model confidence, label toggle. Accessible via internal IP:Port. |
| **Label Export** | CSV/JSON export with `{video_id, frame_timestamps, crops_paths, ground_truth}` |
| **Fine-Tuning** | Head-only fine-tuning on face crops. 10–30 epochs, AdamW, LR=1e-3, early stopping. Runs on
idle GPU slot or off-peak schedule. |
| **Deployment** | Candidate model validated against held-out set → auto-swap if `ΔF1 > 0.02` & confidence
calibration passes. Rollback on regression. |
| **Drift Monitoring** | Weekly confidence histogram comparison. Alerts if `p(C > 0.5)` shifts >10% or review
queue grows unbounded. |
---
## 8. Acceptance Criteria
| Area | Criteria |
|------|----------|
| **Sampling** | Configurable interval honored ±1 frame; handles variable FPS & codec edge cases. |
| **Inference** | Sustained ≥ 30 videos/hour/GPU at default sampling; GPU RAM ≤ 18GB. |
| **Confidence** | Calibration error (ECE) ≤ 0.08 on validation set; thresholds configurable via YAML/ENV. |
| **Routing** | 100% of videos assigned to exactly one bucket; metadata persisted before cleanup. |
| **Active Learning** | Labeled reviews → fine-tune → model swap → next batch uses new weights. Full loop < 48h. |
| **Reliability** | Batch resumes after crash without reprocessing; idempotent file handling. |
| **Syncing** | Scanning process accurately identifies new, modified, and deleted files in the 30TB corpus relative to DB state. |
| **Codec Handling** | Files with unsupported codecs are marked `UNSCANNABLE` in DB without crashing the batch. |
| **Deployment** | `docker-compose.yml` launches DB, Worker(s), and UI. UI accessible on LAN via standard port. |
---
## 9. Implementation Story Mapping
| Epic | User Story | Acceptance Criteria | Priority |
|------|------------|---------------------|----------|
| **E1: Core Pipeline** | As an engineer, I can configure frame sampling interval and extract frames uniformly. |
FFmpeg/OpenCV extracts frames at interval; handles variable duration; logs frame count. | P0 |
| **E1** | As an engineer, I can run face detection on sampled frames in batch. | YOLOv8n/RetinaFace runs at
320²/416²; batch size auto-tuned to VRAM; outputs crop tensors. | P0 |
| **E1** | I can classify face crops and aggregate video confidence. | MobileNetV3 head outputs
calibrated `p`; aggregation logic configurable; results saved to Parquet. | P0 |
| **E1** | As a system, I can detect video resolution and codec. | FFmpeg probe extracts width, height, codec name. Stored in DB. | P0 |
| **E1** | As a system, I can handle unsupported codecs. | If decoding fails, file marked `UNSCANNABLE`. Error logged. | P0 |
| **E2: Routing & Review** | As an operator, I can route videos based on confidence thresholds. | High→MATCH,
Mid→REVIEW, Low→SKIP; thresholds in config; routing logged. | P0 |
| **E2** | As an annotator, I can view low-confidence videos and label them. | UI shows video + frames; label
saved; export triggers training pipeline. | P1 |
| **E3: Active Learning** | As an ML engineer, I can fine-tune the classifier head with reviewed data. | Dataset
versioned; head-only training; early stopping; ECE validated. | P1 |
| **E3** | As a system, I can swap models safely and roll back on regression. | Model registry with
candidate/active states; auto-validate; rollback script. | P1 |
| **E4: Operations** | As a DevOps, I can schedule, monitor, and resume batch jobs. | Multi-GPU process split;
crash recovery; Prometheus metrics; structured logs. | P0 |
| **E4** | As a lead, I can track confidence drift and review queue health. | Dashboard shows distribution, queue
size, throughput; alerts on thresholds. | P1 |
| **E5: Data Management** | As a system, I can scan directories and sync file state to MariaDB. | Detects new/removed files; updates `last_scan_time`; handles 30TB path space efficiently. | P0 |
| **E6: Infrastructure** | As a DevOps, I can deploy the entire stack via Docker Compose. | `docker-compose.yml` includes DB, Worker, UI. GPU passthrough configured. No SSL/Auth. | P0 |
---
## 10. Assumptions & Open Questions
| Item | Status |
|------|--------|
| Dataset size & growth rate | ✅ ~30TB, slow growth (archival). |
| Video codec/resolution distribution | ✅ Varies 480p–4K. Virtually all codecs present. Unsupported ones flagged. |
| Review annotation bandwidth | ❓ Determines acceptable queue backlog & SLA |
| Retraining frequency preference | ❓ Nightly, weekly, or queue-depth triggered? |
| UI deployment environment | ✅ Internal LAN only. No security/SSL needed. Docker Compose managed. |
| Legal/ethics sign-off | ✅ Assumed handled per prompt; document retention policy still needed |
---
## Next Steps
1. **Stakeholder Review**: Validate thresholds, throughput targets, and review workflow.
2. **Sprint 0 Setup**: Provision Docker env, install CUDA 11.8 + PyTorch 2.0.1, set up MariaDB schema & temp storage. Configure GPU passthrough in `docker-compose`.
3. **MVP Build**: Implement E1 stories → baseline throughput & confidence pipeline.
4. **Calibration & Routing**: Add ECE scaling, threshold routing, metadata persistence.
5. **Review & Active Learning**: Deploy UI, wiring label export → fine-tune → model swap.
6. **Data Syncing**: Implement directory scanner (E5) to maintain MariaDB state against 30TB filesystem.
7. **Hardening**: Multi-GPU scheduling, crash recovery, monitoring, drift alerts.
Let me know if you want this exported as a Confluence/Markdown template, or if you'd like detailed technical specs
for any specific story (e.g., FFmpeg sampling logic, TensorRT export pipeline, or Label Studio integration schema).