ποΈ Computer Vision β Dari Pixels ke Scene Understanding
Computer Vision adalah cabang AI yang memungkinkan mesin "melihat" β bukan sekadar menangkap gambar, tapi memahami konten visual. Catatan ini memetakan 6 tingkat hierarki CV dari pixels hingga scene understanding, CNN architecture evolution, object detection (YOLO, DETR), segmentation, dan trade-off matrix. 68 hits di vault β domain ML terbesar ketiga yang belum punya hierarchy.
Daftar Isi
1. Premise β How Machines See
2. Six-Level Hierarchy of Vision
3. Level 1 β Image Acquisition & Preprocessing
4. Level 2 β Feature Extraction
5. Level 3 β Low-Level Vision
6. Level 4 β Mid-Level Vision
7. Level 5 β High-Level Vision
8. Level 6 β Scene Understanding & Reasoning
9. CNN Architecture Timeline
1. Premise β How Machines See
Manusia melihat dengan 200+ juta tahun evolusi β mesin melihat dengan arsitektur yang didesain secara manual :
Pixel Grid (HΓWΓC) β Features β Structures β Objects β Scenes β Meaning
β β β β β β
Raw data Edge, texture Segments Objects Context Decision
Computer Vision itu SUSAH karena:
Ill-posed problem β 3D world diproyeksikan ke 2D β informasi kedalaman hilang
Variability β objek yang sama (kursi) bisa memiliki tampilan tak terbatas
Lighting β pencahayaan mengubah pixel drastis
Scale β objek bisa sangat kecil atau sangat besar di frame
2. Six-Level Hierarchy of Vision
Level Nama Contoh Tugas Input Output L1 Image Acquisition & Preprocessing Capture, denoise, color correction Raw sensor Clean image L2 Feature Extraction Edge detection, SIFT, HOG Image Feature maps L3 Low-Level Vision Segmentation, depth estimation Features Structure maps L4 Mid-Level Vision Object detection, tracking Structure Bounding boxes L5 High-Level Vision Classification, recognition Objects Labels L6 Scene Understanding & Reasoning VQA, captioning, navigation Scene Language/action
3. Level 1 β Image Acquisition & Preprocessing
3.1 Camera Pipeline
Scene β Lens β Sensor (CMOS/CCD) β A/D β Raw β Demosaic β Color β Gamma β JPEG
Sensor types:
Sensor Quantum Efficiency Noise Speed Cost CCD 40-70% Low Slow High CMOS 30-60% Moderate Fast Low Event Camera N/A (delta intensity) Low ΞΌs High
Key:
Resolution (MP) β seberapa detail
Dynamic Range (stops) β rentang gelap ke terang
FPS β seberapa cepat (IoT: 30, ML: 60-1000)
3.2 Preprocessing Operations
Operation Fungsi Algoritma Denoising Hapus noise sensor Gaussian, Median, Bilateral, Non-local Means Deblurring Koreksi blur (motion/focus) Wiener filter, Lucy-Richardson, blind deconvolution Histogram Equalization Perbaiki kontras CLAHE (Contrast Limited Adaptive HE) Gamma Correction Komaasi nonlinear I t e x t o u t β = I t e x t in β t im es g amma White Balance Koreksi warna Gray world, Retinex Resize Ubah dimensi Bilinear, bicubic, lanczos (ML), anti-aliasing
4.1 Classical (Pre-Deep Learning) Features
Feature Type Invariance Application SIFT Keypoint Scale + rotation + illumination Image stitching, 3D reconstruction SURF Keypoint (SIFT + integral image) Scale + rotation Real-time matching ORB Keypoint (FAST + BRIEF) Rotation SLAM, mobile (free) HOG Dense gradient descriptor Illumination Pedestrian detection LBP Texture descriptor Illumination Face recognition Gabor Frequency filter β Texture analysis
4.2 CNN-based Features (Learned)
Layer Depth Feature Type Visualisasi L1 (conv1) Edge, color blobs Lines at various angles L2 (conv2) Textures, patterns Zebra, grid, dots L3 (conv3) Mid-level parts Wheels, eyes, windows L4 (conv4) Object parts Car fronts, faces L5 (conv5) Full objects Cars, dogs, people
5. Level 3 β Low-Level Vision
5.1 Segmentation
Task Definisi Arsitektur Kunci Semantic Segmentation Setiap pixel β class label U-Net, DeepLab, FCN, SegFormer Instance Segmentation Setiap objek β individual mask Mask R-CNN, YOLACT Panoptic Segmentation Stuff (semantic) + things (instance) Panoptic FPN, Mask2Former Part Segmentation Bagian dari objek (wheels, door) PartNet
Benchmark: Cityscapes (50 classes, 5000 images, 1024Γ2048)
Architecture mIoU Cityscapes FPS (T4) DeepLabV3+ (ResNet-101) 82.3% 17 SegFormer-B5 84.0% 11 Mask2Former 85.2% 8 PP-LiteSeg 79.1% 163
5.2 Depth Estimation
Type Output Contoh Monocular Depth map from single image MiDaS, DPT, Depth Anything Stereo Disparity from 2 cameras RAFT-Stereo, NAS-DAD Multi-view Depth from N views COLMAP (SfM)
6. Level 4 β Mid-Level Vision
6.1 Object Detection
Timeline:
R-CNN (2014) β Fast R-CNN (2015) β Faster R-CNN (2015) β YOLO (2016) β SSD (2016)
β RetinaNet (2017) β YOLOv3 (2018) β EfficientDet (2020) β DETR (2020)
β YOLOv8 (2023) β RT-DETR (2023) β YOLOv9 (2024) β YOLOv10 (2024)
6.2 Two-Stage vs One-Stage
Aspek Two-Stage (Faster R-CNN, Mask R-CNN) One-Stage (YOLO, SSD, RetinaNet) Pipeline RPN β RoI β Classify Single shot Accuracy β
Lebih tinggi (mAP+2-5) π‘ Direndahkan (mAP-2-5) Speed 10-30 FPS 60-600+ FPS Trade-off Akurat, lambat Cepat, cukup akurat Use case Autonomous driving, medical Robotics, edge, real-time
6.3 Detection Architecture Comparison
Model Backbone mAP 50-95 FPS (T4) Tahun Cats Faster R-CNN ResNet-50 37.4 15 2015 Two-stage pioneer YOLOv3 DarkNet 33.0 78 2018 One-stage break RetinaNet ResNet-50 36.5 16 2017 Focal Loss EfficientDet-D0 EfficientNet 33.8 98 2020 Efficient YOLOv8m CSPDarknet 50.8 109 2023 Modern versatile RT-DETR-L ResNet-50 53.0 108 2023 Real-time transformer DETR ResNet-50 42.0 12 2020 End-to-end DINO Swin-L 63.2 6 2022 SOTA detection
6.4 Object Tracking
Paradigma Contoh Mekanisme Use Case SORT Kalman filter + IoU matching Fast (260 Hz) DeepSORT SORT + appearance embedding Multi-camera ByteTrack Low + high score boxes Occlusion robust TransTrack Transformer + query High accuracy
7. Level 5 β High-Level Vision
7.1 Image Classification
Evolution:
AlexNet (2012) β VGG (2014) β GoogLeNet (2014) β ResNet (2015)
β DenseNet (2017) β EfficientNet (2019) β ConvNeXt (2022)
AlexNet (2012) β memenangkan ImageNet pertama:
11Γ11 conv, 60M parameters
15.3% top-5 error (vs 26.2% sebelumnya)
GPU training dengan 2Γ GTX 580
ResNet (2015) β skip connections membuat deep network mungkin:
ResNet-50: 25M params, 3.8B FLOPs, 92.1% top-5
ResNet-152: 60M params, 11.3B FLOPs, 93.4% top-5
EfficientNet β compound scaling (depth, width, resolution):
Model Params FLOPs Top-1 EfficientNet-B0 5.3M 0.4B 77.1% EfficientNet-B7 66M 37B 84.3%
ViT (2021) β membawa transformer ke vision:
Image (HΓWΓC) β Patches (PΓP) β Linear β [CLS] token + Position β Transformer encoder β MLP β Class
Model Params ImageNet Top-1 ImageNet Top-1 (22K) ViT-B/16 86M 77.9% 84.1% ViT-L/16 307M 76.5% 85.1% ViT-H/14 632M β 87.8%
Hybrid (CNN + Transformer):
Model Backbone Top-1 FPS DeiT Teacher-student 83.1% Fast Swin-T Shifted windows 81.2% 100+ ConvNeXt Modern CNN 84.3% 80+
8. Level 6 β Scene Understanding & Reasoning
8.1 Visual Question Answering (VQA)
Arsitektur Dataset Accuracy LXMERT VQA v2 72.4% ViLBERT VQA v2 73.3% OFA 12 datasets SOTA multi-task
8.2 Image & Video Captioning
Model Autoregressive Object-Aware CIDEr ViT + GPT2 β
β 120.3 BLIP-2 β
(Q-Former) β
133.2 GIT (Microsoft)β
(Transformer) β
144.1 Flamingo β
(Perceiver) β
β
8.3 Multi-Modal Models (Vision + Language)
Model Vision Encoder Language Decoder Zero-shot? CLIP (OpenAI)ViT-L Text Encoder β
400M pairs BLIP-2 ViT-g OPT/LLaMA β
Instruct LLaVA CLIP ViT-L Vicuna β
SFT GPT-4V β GPT-4 β
RLHF
9. CNN Architecture Timeline + Comparison
Tahun Arsitektur Inovasi Params FLOPs 2012 AlexNet GPU training, ReLU, dropout 60M 0.7B 2014 VGG-16 Small filters (3Γ3), deeper 138M 15.3B 2014 GoogLeNet Inception module 6.8M 1.6B 2015 ResNet-50 Skip connections 25.5M 3.8B 2017 DenseNet-121 Dense connections 8M 2.9B 2018 MobileNetV2 Depthwise separable + inverted residuals 3.5M 0.3B 2019 EfficientNet Compound scaling 5.3-66M 0.4-37B 2021 ViT Pure transformer for vision 86M+ β 2022 ConvNeXt Modernized ResNet 28M 4.5B 2023 SwinV2 Scaled window attention 3B β
10. Cross-Reference ke Vault
References
Szeliski, R. βComputer Vision: Algorithms and Applications.β 2nd ed., Springer, 2022.
Goodfellow, I., Bengio, Y., Courville, A. βDeep Learning.β MIT Press, 2016 β Ch. 9-12.
Krizhevsky, A. et al. βImageNet Classification with Deep Convolutional Neural Networks.β NeurIPS 2012.
He, K. et al. βDeep Residual Learning for Image Recognition.β CVPR 2016.
Ren, S. et al. βFaster R-CNN: Towards Real-Time Object Detection.β NeurIPS 2015.
Redmon, J. & Farhadi, A. βYOLOv3: An Incremental Improvement.β 2018.
Dosovitskiy, A. et al. βAn Image is Worth 16Γ16 Words: Transformers for Image Recognition.β ICLR 2021.
Carion, N. et al. βEnd-to-End Object Detection with Transformers.β ECCV 2020.
Tan, M. & Le, Q. βEfficientNet: Rethinking Model Scaling.β ICML 2019.
Chen, L. et al. βDeepLab: Sematic Image Segmentation.β TPAMI 2018.
Kirillov, A. et al. βSegment Anything.β ICCV 2023.
Radford, A. et al. βLearning Transferable Visual Models From Natural Language Supervision (CLIP).β ICML 2021.
Li, J. et al. βBLIP-2: Bootstrapping Language-Image Pre-training.β 2023.
Liu, Z. et al. βSwin Transformer.β CVPR 2021.