# STORY-09: Observability, Monitoring & Hardening ## Epic **E4: Operations** — As a DevOps engineer, I can schedule, monitor, and resume batch jobs. ## Related Requirements | ID | Requirement | |----|------| | NFR-05 | Fault Tolerance: Auto-retry on transient failures; skip & log on fatal errors | | NFR-06 | Observability: Prometheus/Grafana metrics + structured logging; Tracks FPS, queue depth, confidence distribution, drift alerts | | NFR-01 | Throughput: ≥ 30 videos/hour/GPU | | NFR-03 | GPU Memory Safety: ≤ 18GB per GPU sustained | | NFR-04 | Determinism & Reproducibility: Config-seeded randomness, versioned models | | NFR-07 | Data Volume Handling: Efficient indexing for ~30TB dataset | ## Description Implement comprehensive monitoring, metrics collection, fault tolerance, and system hardening. Expose Prometheus metrics for throughput, GPU utilization, queue depth, and confidence distribution. Build Grafana dashboards for real-time system health. Implement crash recovery, idempotency, retry logic, and drift detection. ## Scope ### In Scope - Prometheus metrics exposure (videos processed, GPU utilization, queue depth, confidence distribution) - Grafana dashboard configuration (throughput, error rates, confidence drift) - Crash recovery with checkpointing - Idempotent processing guarantees - Retry logic for transient errors (up to 3 attempts) - Weekly drift detection job - Alerting on confidence distribution shifts and review queue growth - Worker health checks and auto-restart ### Out of Scope - Frame sampling (covered in STORY-04) - Face detection (covered in STORY-05) - Classification (covered in STORY-06) - Review UI (covered in STORY-08) - Active learning pipeline (covered in STORY-09) ## Deliverables ### 9.1 Prometheus Metrics **File:** `src/metrics.py` Metrics to expose: - **Counter Metrics:** - `videos_processed_total` (label: routing_decision=MATCH|REVIEW|SKIP|UNSCANNABLE|ERROR) - `videos_processed_by_model_total` (label: model_version) - `frames_extracted_total` - `faces_detected_total` - `inference_errors_total` (label: error_type) - `retry_attempts_total` (label: step=probe|extract|detect|classify) - **Gauge Metrics:** - `gpu_utilization_percent` (label: gpu_id) - `gpu_memory_used_bytes` (label: gpu_id) - `gpu_memory_free_bytes` (label: gpu_id) - `queue_depth_pending` - `queue_depth_processing` - `queue_depth_review` - `scratch_usage_bytes` - `scratch_usage_percent` - **Histogram Metrics:** - `video_processing_duration_seconds` (label: routing_decision) - `confidence_score_distribution` (buckets: 0.0, 0.1, 0.2, ..., 0.9, 1.0) - `frame_count_per_video` - `face_count_per_video` - **Summary Metrics:** - `throughput_videos_per_hour` (calculated from counter) - `average_confidence` (calculated from histogram) Implementation: ```python from prometheus_client import start_http_server, Counter, Gauge, Histogram # Start metrics HTTP server start_http_server(9090) # /metrics endpoint # Define metrics videos_processed = Counter( 'videos_processed_total', 'Total videos processed', ['routing_decision'] ) gpu_memory = Gauge( 'gpu_memory_used_bytes', 'GPU memory usage', ['gpu_id'] ) confidence_hist = Histogram( 'confidence_score_distribution', 'Video confidence scores', buckets=[0.0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0] ) ``` ### 9.2 Grafana Dashboards **File:** `monitoring/grafana/dashboards/` Dashboard 1: **System Overview** - Panels: - Throughput (videos/hour) — line chart, 1h window - Queue depth (pending/processing/review) — stacked bar - GPU utilization (both GPUs) — line chart - GPU memory usage (both GPUs) — line chart - Error rate (errors/100 videos) — bar chart - Confidence distribution — histogram Dashboard 2: **Processing Details** - Panels: - Processing duration per video — scatter plot - Frame count per video — histogram - Face count per video — histogram - Confidence score by model version — box plot - Routing decision distribution — pie chart Dashboard 3: **Active Learning** - Panels: - Review queue size over time — line chart - Annotation rate (labels/day) — bar chart - Model version timeline — timeline panel - F1 score by model version — line chart - ECE by model version — line chart **File:** `monitoring/grafana/dashboards/system_overview.json` **File:** `monitoring/grafana/dashboards/processing_details.json` **File:** `monitoring/grafana/dashboards/active_learning.json` ### 9.3 Crash Recovery **File:** `src/crash_recovery.py` Features: - **Checkpointing:** Periodically save processing state ```python def save_checkpoint(video_id, state, progress): checkpoint = { 'video_id': video_id, 'state': state, # 'processing', 'extracting', 'detecting', 'classifying' 'progress': progress, # dict of step -> completed 'timestamp': datetime.utcnow().isoformat() } with open(f'/scratch/checkpoints/{video_id}.json', 'w') as f: json.dump(checkpoint, f) ``` - **Recovery on Startup:** Scan for PROCESSING videos and re-queue them ```sql UPDATE videos SET status = 'PENDING', updated_at = NOW() WHERE status = 'PROCESSING' AND updated_at < NOW() - INTERVAL 5 MINUTE ``` - **Lock Timeout:** 5 minutes (videos stuck in PROCESSING beyond this are re-queued) - **Re-queue Logic:** Only re-queue if worker is down (detected via health check) - **Idempotency:** Re-processing a video produces same results (no duplicates) ### 9.4 Retry Logic **File:** `src/retry.py` Features: - **Retry Decorator:** ```python def retry(max_attempts=3, delay=1.0, backoff=2.0, exceptions=(Exception,)): def decorator(func): def wrapper(*args, **kwargs): last_exception = None for attempt in range(max_attempts): try: return func(*args, **kwargs) except exceptions as e: last_exception = e if attempt < max_attempts - 1: wait = delay * (backoff ** attempt) time.sleep(wait) raise last_exception return wrapper return decorator ``` - **Applicable Steps:** ffprobe, frame extraction, face detection, classification, DB writes - **Transient Errors:** Network timeout, GPU OOM, file lock contention - **Non-Retryable Errors:** Codec unsupported, corrupt file, invalid path → skip and log - **Retry Logging:** Log each retry attempt with error type and delay ### 9.5 Drift Detection **File:** `src/drift_detector.py` Features: - **Weekly Job:** Compare current confidence distribution to baseline ```python def detect_drift(current_confidences, baseline_confidences, threshold=0.10): # Compare p(C > 0.5) shift current_high = sum(1 for c in current_confidences if c > 0.5) / len(current_confidences) baseline_high = sum(1 for c in baseline_confidences if c > 0.5) / len(baseline_confidences) shift = abs(current_high - baseline_high) if shift > threshold: alert(f"Confidence drift detected: {shift:.2%} shift in p(C > 0.5)") return True return False ``` - **Baseline:** Stored in DB or config (computed from last training cycle) - **Alert Conditions:** - `p(C > 0.5)` shifts > 10% from baseline - Review queue grows unbounded (> 1000 items for > 24 hours) - Throughput drops below 20 videos/hour/GPU for > 1 hour - Error rate exceeds 5% for any 1-hour window - **Alert Channels:** Email, Slack webhook, or log entry (configurable) ### 9.6 Worker Health Checks **File:** `src/health_check.py` Features: - **Health Endpoint:** `/health` returns worker status ```json { "status": "healthy", "gpu_available": true, "gpu_memory_used_gb": 12.5, "queue_depth": 42, "uptime_seconds": 86400, "videos_processed_today": 156, "last_error": null } ``` - **Auto-Restart:** Docker restart policy for worker container ```yaml restart: unless-stopped healthcheck: test: ["CMD", "curl", "-f", "http://localhost:9090/health"] interval: 30s timeout: 10s retries: 3 ``` - **GPU Health:** Monitor GPU temperature and error counts - **Disk Health:** Monitor scratch space usage and NAS connectivity ### 9.7 Docker Compose Update **File:** `docker-compose.yml` (update) Add monitoring services: ```yaml prometheus: image: prom/prometheus:v2.48.0 volumes: - ./monitoring/prometheus/prometheus.yml:/etc/prometheus/prometheus.yml - prometheus_data:/prometheus ports: - "9090:9090" networks: - videodetect-network grafana: image: grafana/grafana:10.2.0 volumes: - grafana_data:/var/lib/grafana - ./monitoring/grafana/dashboards:/etc/grafana/provisioning/dashboards environment: - GF_SECURITY_ADMIN_PASSWORD=${GRAFANA_PASSWORD} ports: - "3000:3000" networks: - videodetect-network volumes: prometheus_data: grafana_data: ``` ### 9.8 Configuration Updates **File:** `config.yaml` (updates) New fields: ```yaml monitoring: prometheus: enabled: true port: 9090 metrics_path: /metrics grafana: enabled: true port: 3000 alerts: confidence_drift_threshold: 0.10 review_queue_max_size: 1000 review_queue_max_age_hours: 24 throughput_min_videos_per_hour: 20 throughput_min_duration_hours: 1 error_rate_threshold: 0.05 error_rate_window_hours: 1 drift_detection: enabled: true schedule: weekly # cron: 0 2 * * 0 (Sundays at 2 AM) baseline_source: db # db or config crash_recovery: lock_timeout_minutes: 5 auto_requeue: true retry: max_attempts: 3 initial_delay: 1.0 backoff_factor: 2.0 retryable_errors: - timeout - gpu_oom - file_lock non_retryable_errors: - codec_unsupported - file_corrupt - invalid_path ``` ## Acceptance Criteria ### Functional - [ ] Prometheus metrics are exposed at `/metrics` endpoint and queryable - [ ] All required metrics are present (counters, gauges, histograms, summaries) - [ ] Grafana dashboards load and display correct data - [ ] System Overview dashboard shows throughput, queue depth, GPU metrics - [ ] Processing Details dashboard shows duration, frame count, confidence distribution - [ ] Active Learning dashboard shows review queue, annotation rate, model versions - [ ] Crash recovery re-queues stuck PROCESSING videos on worker restart - [ ] Idempotent processing: re-processing produces same results (no duplicates) - [ ] Retry logic retries transient errors up to 3 times with exponential backoff - [ ] Non-retryable errors are skipped and logged (no infinite retry) - [ ] Drift detection runs weekly and alerts on > 10% confidence shift - [ ] Review queue growth alert triggers when queue > 1000 for > 24 hours - [ ] Worker health check returns correct status - [ ] Auto-restart triggers on health check failure ### Non-Functional - [ ] Metrics collection overhead < 2% of CPU - [ ] Grafana dashboard loads in < 3 seconds - [ ] Crash recovery completes in < 30 seconds - [ ] Drift detection completes in < 5 minutes - [ ] Alert delivery completes in < 60 seconds - [ ] Prometheus data retention: 30 days (configurable) - [ ] Grafana data source refresh: 30 seconds ### Technical Constraints - [ ] Prometheus metrics follow naming conventions (unit suffixes, proper labels) - [ ] Grafana dashboards are JSON-exportable and version-controlled - [ ] Crash recovery is idempotent (running twice produces same result) - [ ] Retry logic does not retry non-transient errors - [ ] Drift detection baseline is stored and versioned - [ ] Health check endpoint responds in < 1 second - [ ] All alerts are logged with timestamp and context ## Dependencies - **Prerequisites:** STORY-01 (Foundation), STORY-05 (Classification — provides metrics data), STORY-08 (Active Learning — provides model metrics) - **Depends on:** None (can be built in parallel with other stories) - **Enables:** Production deployment and long-term operation ## Risks & Mitigations | Risk | Mitigation | |------|--| | Prometheus metrics cardinality explosion | Limit label cardinality; use histograms instead of individual values | | Grafana dashboard load time with large datasets | Use Prometheus aggregations; pre-compute panels | | Crash recovery misses in-flight writes | Use DB transactions; lock timeout prevents permanent locks | | Drift detection baseline becomes stale | Update baseline with each training cycle | | Alert fatigue from too many alerts | Tune thresholds; implement alert grouping | ## Estimated Effort - **Sprint:** 9+ - **Story Points:** 34 - **Dependencies:** STORY-01, STORY-05, STORY-08