7.6 KiB
7.6 KiB
STORY-01: Foundation & Infrastructure
Epic
E6: Infrastructure — As a DevOps engineer, I can deploy the entire stack via Docker Compose.
Related Requirements
| ID | Requirement |
|---|---|
| TC-01 | Hardware: 2× Tesla P40 24GB (compute capability 5.2, PCIe 3.0, no Tensor Cores) |
| TC-02 | CUDA/Torch Compatibility: CUDA ≤ 11.8, PyTorch ≤ 2.1.0, FP32 inference only |
| TC-03 | Storage I/O: Fast local NVMe/SSD for temp frame cache; shared NAS/SMB for video input/output |
| TC-04 | Framework Stack: PyTorch → ONNX → TensorRT FP32; FFmpeg/OpenCV; MariaDB |
| TC-05 | Deployment Model: Docker Compose orchestrates all services; GPUs passed via nvidia-container-toolkit |
| TC-06 | Network Security: Internal LAN only; no reverse proxy, SSL, or auth |
| NFR-04 | Determinism & Reproducibility: Config-seeded randomness, versioned models |
Description
Establish the Docker environment, database schema, and basic connectivity. No video processing logic yet — just the operational skeleton that all subsequent stories depend on.
Scope
In Scope
- Docker Compose multi-service orchestration
- Custom Worker Dockerfile with CUDA 11.8 + PyTorch 2.1.0 + TensorRT FP32
- MariaDB schema design and migration scripts
- Volume mounts for NVMe scratch, NAS input/output, and model persistence
- Configuration management (config.yaml + environment variables)
- Structured logging setup (JSON format to stdout)
- Database connection layer with connection pooling
- Network connectivity verification between all services
Out of Scope
- Video processing logic (covered in STORY-02 through STORY-05)
- Model loading or inference (covered in STORY-04)
- Review UI functionality (covered in STORY-06)
- Monitoring dashboards (covered in STORY-08)
- Active learning pipeline (covered in STORY-07)
Deliverables
1.1 Docker Compose Structure
File: docker-compose.yml
Services defined:
mariadb: MariaDB 10.11+ with persistent volumeworker: PyTorch/TensorRT inference worker with GPU passthroughui: Placeholder service (nginx serving static page) for network verification
Key configurations:
nvidia-container-toolkitruntime configuration for GPU passthrough- Volume mounts:
- NVMe →
/scratch(tmpfs for speed) - NAS/SMB →
/data/inputand/data/output - Persistent →
/modelsand/data/training
- NVMe →
- Network bridge for inter-service communication
- Resource limits (GPU memory caps per NFR-03)
1.2 Worker Dockerfile
File: worker/Dockerfile
Base image: nvidia/cuda:11.8.0-runtime-ubuntu22.04
Installed packages:
- PyTorch 2.1.0 (CUDA 11.8, FP32 only)
- TensorRT 8.6+ (FP32)
- OpenCV 4.8+
- FFmpeg 5.x + ffprobe
- ONNX Runtime
- Python 3.10+
- Required system libraries (libcudnn8, libglib2.0, etc.)
1.3 Database Schema
File: db/schema.sql
Tables:
videos:id(BIGINT PK),file_path(VARCHAR),file_hash(CHAR(64)),resolution_w(INT),resolution_h(INT),codec(VARCHAR),duration(FLOAT),status(ENUM: NEW, PENDING, PROCESSING, COMPLETED, UNSCANNABLE, ERROR),last_scan_time(DATETIME),last_processed_time(DATETIME),created_at(DATETIME),updated_at(DATETIME)processing_logs:id(BIGINT PK),video_id(BIGINT FK),model_version(VARCHAR),frame_count(INT),confidence_score(FLOAT),routing_decision(ENUM: MATCH, REVIEW, SKIP),processed_at(DATETIME),error_message(TEXT)models:version(VARCHAR PK),status(ENUM: ACTIVE, CANDIDATE, ARCHIVED),path(VARCHAR),calibration_temp(FLOAT),f1_score(FLOAT),ece_score(FLOAT),deployed_at(DATETIME),created_at(DATETIME)review_queue:id(BIGINT PK),video_id(BIGINT FK),confidence_score(FLOAT),routing_decision(ENUM: REVIEW),annotated(BOOLEAN),ground_truth(BOOLEAN),annotated_at(DATETIME),created_at(DATETIME)
Indexes:
idx_videos_statusonvideos(status)idx_videos_file_hashonvideos(file_hash)(UNIQUE)idx_videos_last_scanonvideos(last_scan_time)idx_processing_logs_videoonprocessing_logs(video_id)idx_models_statusonmodels(status)idx_review_queue_annotatedonreview_queue(annotated)
1.4 Database Connection Layer
File: src/db_connector.py
Features:
- Connection pooling (DBUtils + PyMySQL)
- Configurable pool size (min=5, max=20)
- Automatic reconnection on disconnect
- Context manager support
- Prepared statements for all queries
- Transaction support for atomic state transitions
1.5 Configuration Management
File: config.yaml
Contents:
sampling:interval_seconds: 30,override_per_job: truethresholds:T_high: 0.75,T_low: 0.45gpu:max_memory_gb: 18,batch_size: auto,device: cudastorage:scratch_path: /scratch,input_path: /data/input,output_path: /data/outputdatabase:host: mariadb,port: 3306,pool_size: 20logging:format: json,level: INFOmodel:face_detector: yolo8n,classifier: mobilenetv3,input_size: 224
File: src/config_loader.py
Features:
- Load config.yaml with environment variable overrides
- Validate all required fields
- Provide typed accessors (e.g.,
config.thresholds.T_high) - Hot-reload support for config changes
1.6 Logging Setup
File: src/logging_config.py
- JSON structured logging via
python-json-logger - Log levels: DEBUG, INFO, WARNING, ERROR, CRITICAL
- Fields: timestamp, level, service, video_id, message, metadata (key-value pairs)
- Log rotation: 100MB per file, 10 files max
- All logs to stdout for Docker capture
Acceptance Criteria
Functional
docker-compose upstarts MariaDB, Worker, and UI containers successfully- Worker container can connect to MariaDB and execute schema migrations
- GPU is visible inside Worker container (
nvidia-smishows Tesla P40) - Volume mounts are accessible and writable in all containers
- Config.yaml loads correctly with all required fields validated
- Structured logging produces valid JSON output in all services
- Database connection pool handles concurrent connections (test with 20 simultaneous)
Non-Functional
- Worker container starts within 30 seconds
- MariaDB starts within 15 seconds
- GPU memory usage in Worker is < 2GB at idle (before model loading)
- All services communicate over Docker internal network (no host network exposure except UI port)
- Schema migration is idempotent (running twice produces same result)
Technical Constraints
- CUDA version in Worker is 11.8 (verified via
torch.version.cuda) - PyTorch version ≤ 2.1.0 (verified via
torch.__version__) - TensorRT runs in FP32 mode only
- No Tensor Cores used (compute capability 5.2 constraint respected)
- No SSL, auth, or reverse proxy configured (TC-06)
Dependencies
- Prerequisites: NVIDIA Container Toolkit installed on host, Docker Compose v2+, NAS/SMB mounts configured
- Depends on: None (this is the foundational story)
- Enables: STORY-02 (Ingestion), STORY-03 (Orchestration), STORY-04 (Inference), STORY-05 (Results), STORY-06 (Review UI), STORY-07 (Active Learning), STORY-08 (Monitoring)
Risks & Mitigations
| Risk | Mitigation |
|---|---|
| Tesla P40 (CC 5.2) incompatible with newer TensorRT | Use TensorRT 8.6 which supports CC 5.x; test early |
| CUDA 11.8 + PyTorch 2.1.0 dependency conflicts | Pin all versions in Dockerfile; use nvidia base image |
| NAS/SMB mount latency affects processing | Use local tmpfs for scratch; only read from NAS |
| MariaDB connection pool exhaustion | Monitor pool metrics; tune pool_size based on worker count |
Estimated Effort
- Sprint: 1-2
- Story Points: 13
- Dependencies: None