171 lines
7.6 KiB
Markdown
171 lines
7.6 KiB
Markdown
# STORY-01: Foundation & Infrastructure
|
||
|
||
## Epic
|
||
**E6: Infrastructure** — As a DevOps engineer, I can deploy the entire stack via Docker Compose.
|
||
|
||
## Related Requirements
|
||
| ID | Requirement |
|
||
|----|-------------|
|
||
| TC-01 | Hardware: 2× Tesla P40 24GB (compute capability 5.2, PCIe 3.0, no Tensor Cores) |
|
||
| TC-02 | CUDA/Torch Compatibility: CUDA ≤ 11.8, PyTorch ≤ 2.1.0, FP32 inference only |
|
||
| TC-03 | Storage I/O: Fast local NVMe/SSD for temp frame cache; shared NAS/SMB for video input/output |
|
||
| TC-04 | Framework Stack: PyTorch → ONNX → TensorRT FP32; FFmpeg/OpenCV; MariaDB |
|
||
| TC-05 | Deployment Model: Docker Compose orchestrates all services; GPUs passed via nvidia-container-toolkit |
|
||
| TC-06 | Network Security: Internal LAN only; no reverse proxy, SSL, or auth |
|
||
| NFR-04 | Determinism & Reproducibility: Config-seeded randomness, versioned models |
|
||
|
||
## Description
|
||
Establish the Docker environment, database schema, and basic connectivity. No video processing logic yet — just the operational skeleton that all subsequent stories depend on.
|
||
|
||
## Scope
|
||
|
||
### In Scope
|
||
- Docker Compose multi-service orchestration
|
||
- Custom Worker Dockerfile with CUDA 11.8 + PyTorch 2.1.0 + TensorRT FP32
|
||
- MariaDB schema design and migration scripts
|
||
- Volume mounts for NVMe scratch, NAS input/output, and model persistence
|
||
- Configuration management (config.yaml + environment variables)
|
||
- Structured logging setup (JSON format to stdout)
|
||
- Database connection layer with connection pooling
|
||
- Network connectivity verification between all services
|
||
|
||
### Out of Scope
|
||
- Video processing logic (covered in STORY-02 through STORY-05)
|
||
- Model loading or inference (covered in STORY-04)
|
||
- Review UI functionality (covered in STORY-06)
|
||
- Monitoring dashboards (covered in STORY-08)
|
||
- Active learning pipeline (covered in STORY-07)
|
||
|
||
## Deliverables
|
||
|
||
### 1.1 Docker Compose Structure
|
||
**File:** `docker-compose.yml`
|
||
|
||
Services defined:
|
||
- `mariadb`: MariaDB 10.11+ with persistent volume
|
||
- `worker`: PyTorch/TensorRT inference worker with GPU passthrough
|
||
- `ui`: Placeholder service (nginx serving static page) for network verification
|
||
|
||
Key configurations:
|
||
- `nvidia-container-toolkit` runtime configuration for GPU passthrough
|
||
- Volume mounts:
|
||
- NVMe → `/scratch` (tmpfs for speed)
|
||
- NAS/SMB → `/data/input` and `/data/output`
|
||
- Persistent → `/models` and `/data/training`
|
||
- Network bridge for inter-service communication
|
||
- Resource limits (GPU memory caps per NFR-03)
|
||
|
||
### 1.2 Worker Dockerfile
|
||
**File:** `worker/Dockerfile`
|
||
|
||
Base image: `nvidia/cuda:11.8.0-runtime-ubuntu22.04`
|
||
|
||
Installed packages:
|
||
- PyTorch 2.1.0 (CUDA 11.8, FP32 only)
|
||
- TensorRT 8.6+ (FP32)
|
||
- OpenCV 4.8+
|
||
- FFmpeg 5.x + ffprobe
|
||
- ONNX Runtime
|
||
- Python 3.10+
|
||
- Required system libraries (libcudnn8, libglib2.0, etc.)
|
||
|
||
### 1.3 Database Schema
|
||
**File:** `db/schema.sql`
|
||
|
||
Tables:
|
||
- `videos`: `id` (BIGINT PK), `file_path` (VARCHAR), `file_hash` (CHAR(64)), `resolution_w` (INT), `resolution_h` (INT), `codec` (VARCHAR), `duration` (FLOAT), `status` (ENUM: NEW, PENDING, PROCESSING, COMPLETED, UNSCANNABLE, ERROR), `last_scan_time` (DATETIME), `last_processed_time` (DATETIME), `created_at` (DATETIME), `updated_at` (DATETIME)
|
||
- `processing_logs`: `id` (BIGINT PK), `video_id` (BIGINT FK), `model_version` (VARCHAR), `frame_count` (INT), `confidence_score` (FLOAT), `routing_decision` (ENUM: MATCH, REVIEW, SKIP), `processed_at` (DATETIME), `error_message` (TEXT)
|
||
- `models`: `version` (VARCHAR PK), `status` (ENUM: ACTIVE, CANDIDATE, ARCHIVED), `path` (VARCHAR), `calibration_temp` (FLOAT), `f1_score` (FLOAT), `ece_score` (FLOAT), `deployed_at` (DATETIME), `created_at` (DATETIME)
|
||
- `review_queue`: `id` (BIGINT PK), `video_id` (BIGINT FK), `confidence_score` (FLOAT), `routing_decision` (ENUM: REVIEW), `annotated` (BOOLEAN), `ground_truth` (BOOLEAN), `annotated_at` (DATETIME), `created_at` (DATETIME)
|
||
|
||
Indexes:
|
||
- `idx_videos_status` on `videos(status)`
|
||
- `idx_videos_file_hash` on `videos(file_hash)` (UNIQUE)
|
||
- `idx_videos_last_scan` on `videos(last_scan_time)`
|
||
- `idx_processing_logs_video` on `processing_logs(video_id)`
|
||
- `idx_models_status` on `models(status)`
|
||
- `idx_review_queue_annotated` on `review_queue(annotated)`
|
||
|
||
### 1.4 Database Connection Layer
|
||
**File:** `src/db_connector.py`
|
||
|
||
Features:
|
||
- Connection pooling (DBUtils + PyMySQL)
|
||
- Configurable pool size (min=5, max=20)
|
||
- Automatic reconnection on disconnect
|
||
- Context manager support
|
||
- Prepared statements for all queries
|
||
- Transaction support for atomic state transitions
|
||
|
||
### 1.5 Configuration Management
|
||
**File:** `config.yaml`
|
||
|
||
Contents:
|
||
- `sampling`: `interval_seconds: 30`, `override_per_job: true`
|
||
- `thresholds`: `T_high: 0.75`, `T_low: 0.45`
|
||
- `gpu`: `max_memory_gb: 18`, `batch_size: auto`, `device: cuda`
|
||
- `storage`: `scratch_path: /scratch`, `input_path: /data/input`, `output_path: /data/output`
|
||
- `database`: `host: mariadb`, `port: 3306`, `pool_size: 20`
|
||
- `logging`: `format: json`, `level: INFO`
|
||
- `model`: `face_detector: yolo8n`, `classifier: mobilenetv3`, `input_size: 224`
|
||
|
||
**File:** `src/config_loader.py`
|
||
|
||
Features:
|
||
- Load config.yaml with environment variable overrides
|
||
- Validate all required fields
|
||
- Provide typed accessors (e.g., `config.thresholds.T_high`)
|
||
- Hot-reload support for config changes
|
||
|
||
### 1.6 Logging Setup
|
||
**File:** `src/logging_config.py`
|
||
|
||
- JSON structured logging via `python-json-logger`
|
||
- Log levels: DEBUG, INFO, WARNING, ERROR, CRITICAL
|
||
- Fields: timestamp, level, service, video_id, message, metadata (key-value pairs)
|
||
- Log rotation: 100MB per file, 10 files max
|
||
- All logs to stdout for Docker capture
|
||
|
||
## Acceptance Criteria
|
||
|
||
### Functional
|
||
- [ ] `docker-compose up` starts MariaDB, Worker, and UI containers successfully
|
||
- [ ] Worker container can connect to MariaDB and execute schema migrations
|
||
- [ ] GPU is visible inside Worker container (`nvidia-smi` shows Tesla P40)
|
||
- [ ] Volume mounts are accessible and writable in all containers
|
||
- [ ] Config.yaml loads correctly with all required fields validated
|
||
- [ ] Structured logging produces valid JSON output in all services
|
||
- [ ] Database connection pool handles concurrent connections (test with 20 simultaneous)
|
||
|
||
### Non-Functional
|
||
- [ ] Worker container starts within 30 seconds
|
||
- [ ] MariaDB starts within 15 seconds
|
||
- [ ] GPU memory usage in Worker is < 2GB at idle (before model loading)
|
||
- [ ] All services communicate over Docker internal network (no host network exposure except UI port)
|
||
- [ ] Schema migration is idempotent (running twice produces same result)
|
||
|
||
### Technical Constraints
|
||
- [ ] CUDA version in Worker is 11.8 (verified via `torch.version.cuda`)
|
||
- [ ] PyTorch version ≤ 2.1.0 (verified via `torch.__version__`)
|
||
- [ ] TensorRT runs in FP32 mode only
|
||
- [ ] No Tensor Cores used (compute capability 5.2 constraint respected)
|
||
- [ ] No SSL, auth, or reverse proxy configured (TC-06)
|
||
|
||
## Dependencies
|
||
- **Prerequisites:** NVIDIA Container Toolkit installed on host, Docker Compose v2+, NAS/SMB mounts configured
|
||
- **Depends on:** None (this is the foundational story)
|
||
- **Enables:** STORY-02 (Ingestion), STORY-03 (Orchestration), STORY-04 (Inference), STORY-05 (Results), STORY-06 (Review UI), STORY-07 (Active Learning), STORY-08 (Monitoring)
|
||
|
||
## Risks & Mitigations
|
||
| Risk | Mitigation |
|
||
|------|-----------|
|
||
| Tesla P40 (CC 5.2) incompatible with newer TensorRT | Use TensorRT 8.6 which supports CC 5.x; test early |
|
||
| CUDA 11.8 + PyTorch 2.1.0 dependency conflicts | Pin all versions in Dockerfile; use nvidia base image |
|
||
| NAS/SMB mount latency affects processing | Use local tmpfs for scratch; only read from NAS |
|
||
| MariaDB connection pool exhaustion | Monitor pool metrics; tune pool_size based on worker count |
|
||
|
||
## Estimated Effort
|
||
- **Sprint:** 1-2
|
||
- **Story Points:** 13
|
||
- **Dependencies:** None
|