# STORY-01: Foundation & Infrastructure ## Epic **E6: Infrastructure** — As a DevOps engineer, I can deploy the entire stack via Docker Compose. ## Related Requirements | ID | Requirement | |----|-------------| | TC-01 | Hardware: 2× Tesla P40 24GB (compute capability 5.2, PCIe 3.0, no Tensor Cores) | | TC-02 | CUDA/Torch Compatibility: CUDA ≤ 11.8, PyTorch ≤ 2.1.0, FP32 inference only | | TC-03 | Storage I/O: Fast local NVMe/SSD for temp frame cache; shared NAS/SMB for video input/output | | TC-04 | Framework Stack: PyTorch → ONNX → TensorRT FP32; FFmpeg/OpenCV; MariaDB | | TC-05 | Deployment Model: Docker Compose orchestrates all services; GPUs passed via nvidia-container-toolkit | | TC-06 | Network Security: Internal LAN only; no reverse proxy, SSL, or auth | | NFR-04 | Determinism & Reproducibility: Config-seeded randomness, versioned models | ## Description Establish the Docker environment, database schema, and basic connectivity. No video processing logic yet — just the operational skeleton that all subsequent stories depend on. ## Scope ### In Scope - Docker Compose multi-service orchestration - Custom Worker Dockerfile with CUDA 11.8 + PyTorch 2.1.0 + TensorRT FP32 - MariaDB schema design and migration scripts - Volume mounts for NVMe scratch, NAS input/output, and model persistence - Configuration management (config.yaml + environment variables) - Structured logging setup (JSON format to stdout) - Database connection layer with connection pooling - Network connectivity verification between all services ### Out of Scope - Video processing logic (covered in STORY-02 through STORY-05) - Model loading or inference (covered in STORY-04) - Review UI functionality (covered in STORY-06) - Monitoring dashboards (covered in STORY-08) - Active learning pipeline (covered in STORY-07) ## Deliverables ### 1.1 Docker Compose Structure **File:** `docker-compose.yml` Services defined: - `mariadb`: MariaDB 10.11+ with persistent volume - `worker`: PyTorch/TensorRT inference worker with GPU passthrough - `ui`: Placeholder service (nginx serving static page) for network verification Key configurations: - `nvidia-container-toolkit` runtime configuration for GPU passthrough - Volume mounts: - NVMe → `/scratch` (tmpfs for speed) - NAS/SMB → `/data/input` and `/data/output` - Persistent → `/models` and `/data/training` - Network bridge for inter-service communication - Resource limits (GPU memory caps per NFR-03) ### 1.2 Worker Dockerfile **File:** `worker/Dockerfile` Base image: `nvidia/cuda:11.8.0-runtime-ubuntu22.04` Installed packages: - PyTorch 2.1.0 (CUDA 11.8, FP32 only) - TensorRT 8.6+ (FP32) - OpenCV 4.8+ - FFmpeg 5.x + ffprobe - ONNX Runtime - Python 3.10+ - Required system libraries (libcudnn8, libglib2.0, etc.) ### 1.3 Database Schema **File:** `db/schema.sql` Tables: - `videos`: `id` (BIGINT PK), `file_path` (VARCHAR), `file_hash` (CHAR(64)), `resolution_w` (INT), `resolution_h` (INT), `codec` (VARCHAR), `duration` (FLOAT), `status` (ENUM: NEW, PENDING, PROCESSING, COMPLETED, UNSCANNABLE, ERROR), `last_scan_time` (DATETIME), `last_processed_time` (DATETIME), `created_at` (DATETIME), `updated_at` (DATETIME) - `processing_logs`: `id` (BIGINT PK), `video_id` (BIGINT FK), `model_version` (VARCHAR), `frame_count` (INT), `confidence_score` (FLOAT), `routing_decision` (ENUM: MATCH, REVIEW, SKIP), `processed_at` (DATETIME), `error_message` (TEXT) - `models`: `version` (VARCHAR PK), `status` (ENUM: ACTIVE, CANDIDATE, ARCHIVED), `path` (VARCHAR), `calibration_temp` (FLOAT), `f1_score` (FLOAT), `ece_score` (FLOAT), `deployed_at` (DATETIME), `created_at` (DATETIME) - `review_queue`: `id` (BIGINT PK), `video_id` (BIGINT FK), `confidence_score` (FLOAT), `routing_decision` (ENUM: REVIEW), `annotated` (BOOLEAN), `ground_truth` (BOOLEAN), `annotated_at` (DATETIME), `created_at` (DATETIME) Indexes: - `idx_videos_status` on `videos(status)` - `idx_videos_file_hash` on `videos(file_hash)` (UNIQUE) - `idx_videos_last_scan` on `videos(last_scan_time)` - `idx_processing_logs_video` on `processing_logs(video_id)` - `idx_models_status` on `models(status)` - `idx_review_queue_annotated` on `review_queue(annotated)` ### 1.4 Database Connection Layer **File:** `src/db_connector.py` Features: - Connection pooling (DBUtils + PyMySQL) - Configurable pool size (min=5, max=20) - Automatic reconnection on disconnect - Context manager support - Prepared statements for all queries - Transaction support for atomic state transitions ### 1.5 Configuration Management **File:** `config.yaml` Contents: - `sampling`: `interval_seconds: 30`, `override_per_job: true` - `thresholds`: `T_high: 0.75`, `T_low: 0.45` - `gpu`: `max_memory_gb: 18`, `batch_size: auto`, `device: cuda` - `storage`: `scratch_path: /scratch`, `input_path: /data/input`, `output_path: /data/output` - `database`: `host: mariadb`, `port: 3306`, `pool_size: 20` - `logging`: `format: json`, `level: INFO` - `model`: `face_detector: yolo8n`, `classifier: mobilenetv3`, `input_size: 224` **File:** `src/config_loader.py` Features: - Load config.yaml with environment variable overrides - Validate all required fields - Provide typed accessors (e.g., `config.thresholds.T_high`) - Hot-reload support for config changes ### 1.6 Logging Setup **File:** `src/logging_config.py` - JSON structured logging via `python-json-logger` - Log levels: DEBUG, INFO, WARNING, ERROR, CRITICAL - Fields: timestamp, level, service, video_id, message, metadata (key-value pairs) - Log rotation: 100MB per file, 10 files max - All logs to stdout for Docker capture ## Acceptance Criteria ### Functional - [ ] `docker-compose up` starts MariaDB, Worker, and UI containers successfully - [ ] Worker container can connect to MariaDB and execute schema migrations - [ ] GPU is visible inside Worker container (`nvidia-smi` shows Tesla P40) - [ ] Volume mounts are accessible and writable in all containers - [ ] Config.yaml loads correctly with all required fields validated - [ ] Structured logging produces valid JSON output in all services - [ ] Database connection pool handles concurrent connections (test with 20 simultaneous) ### Non-Functional - [ ] Worker container starts within 30 seconds - [ ] MariaDB starts within 15 seconds - [ ] GPU memory usage in Worker is < 2GB at idle (before model loading) - [ ] All services communicate over Docker internal network (no host network exposure except UI port) - [ ] Schema migration is idempotent (running twice produces same result) ### Technical Constraints - [ ] CUDA version in Worker is 11.8 (verified via `torch.version.cuda`) - [ ] PyTorch version ≤ 2.1.0 (verified via `torch.__version__`) - [ ] TensorRT runs in FP32 mode only - [ ] No Tensor Cores used (compute capability 5.2 constraint respected) - [ ] No SSL, auth, or reverse proxy configured (TC-06) ## Dependencies - **Prerequisites:** NVIDIA Container Toolkit installed on host, Docker Compose v2+, NAS/SMB mounts configured - **Depends on:** None (this is the foundational story) - **Enables:** STORY-02 (Ingestion), STORY-03 (Orchestration), STORY-04 (Inference), STORY-05 (Results), STORY-06 (Review UI), STORY-07 (Active Learning), STORY-08 (Monitoring) ## Risks & Mitigations | Risk | Mitigation | |------|-----------| | Tesla P40 (CC 5.2) incompatible with newer TensorRT | Use TensorRT 8.6 which supports CC 5.x; test early | | CUDA 11.8 + PyTorch 2.1.0 dependency conflicts | Pin all versions in Dockerfile; use nvidia base image | | NAS/SMB mount latency affects processing | Use local tmpfs for scratch; only read from NAS | | MariaDB connection pool exhaustion | Monitor pool metrics; tune pool_size based on worker count | ## Estimated Effort - **Sprint:** 1-2 - **Story Points:** 13 - **Dependencies:** None