🎮 GPU Programming & Parallel Compute — Deep Dive: CUDA, ROCm, Vulkan Compute, GPU Architecture, CUDA Cores vs Tensor Cores
Panduan komprehensif GPU sebagai compute engine — dari arsitektur hardware (SM, warp, memory hierarchy) sampai programming model (CUDA, ROCm, Vulkan Compute, OpenCL, SYCL). Mencakup GPU architecture (Ampere, Hopper, RDNA3), memory hierarchy (global, shared, local, constant, texture), parallel programming patterns (grid-stride loop, reduction, scan, tiling), CUDA/ROCm programming model (kernel, grid, block, thread, shared memory), GPU-accelerated ML training (mixed precision, tensor cores, distributed training), GPU compute untuk non-ML (hashcat, password cracking, signal processing, rendering), dan perbandingan platform GPU (NVIDIA CUDA vs AMD ROCm vs Intel oneAPI vs Apple Metal). Vault punya hierarchy-classical-ml-algorithms dan berbagai AI notes yang bergantung pada GPU — catatan ini adalah fondasi hardware compute-nya.
Data Parallel: tiap GPU punya model copy penuh, batch dibagi
→ Gradient all-reduce setelah tiap step
Tensor Parallel: 1 layer dibagi ke beberapa GPU
→ Komunikasi setiap forward/backward
Pipeline Parallel: layer dibagi, tiap GPU pegang contiguous layers
→ Micro-batching untuk overlap compute + communication
Framework Comparison
Framework
CUDA Support
ROCm Support
Mixed Precision
Distributed
Best For
PyTorch
✅ Native
✅ 5.7+
✅ AMP + FSDP
✅ DDP/FSDP/HF
Research, production
TensorFlow
✅ Native
✅ (limited)
✅ Mixed precision
✅ Mirrored + PS
Production pipeline
JAX
✅ Native
❌ No
✅ Native
✅ pmap + pjit
Research, performance
MosaicML Composer
✅
❌
✅
✅ FSDP
Training efficiency
GPU untuk Non-ML Compute
Aplikasi
Tool
GPU Speedup vs CPU
Notes
Password Cracking
hashcat
100-1000x
MD5: ~100 GH/s on RTX 4090
Signal Processing
cuFFT, GNU Radio
10-100x
FFT, FIR filter, convolution
Video Encoding
NVENC, FFmpeg
10-50x
H.264/H.265 hardware encoder
Ray Tracing
OptiX, Vulkan RT
Real-time
Global illumination, caustics
Scientific Compute
cuBLAS, cuSOLVER
10-100x
Linear algebra, sparse solvers
Database Acceleration
HeavyDB, RAPIDS
10-50x
SQL query GPU-accelerated
Genomics
GATK, Parabricks
10-50x
DNA sequencing alignment
Platform Comparison
Aspek
NVIDIA CUDA
AMD ROCm
Intel oneAPI
Apple Metal
Hardware
GeForce, Quadro, Tesla
Radeon, Instinct
Arc, Flex, Max
Apple Silicon
Programming
CUDA C++
HIP C++
SYCL, DPC++
Metal Shading Language
ML Framework
PyTorch, TF, JAX
PyTorch (5.7+), TF (limited)
PyTorch (oneDNN)
CoreML, MPS
CUDA Compatibility
✅ Native
⚠️ HIP porting layer
❌ No
❌ No
Ecosystem Maturity
Sangat matang
Growing
Limited
Mature for Apple
Best For
ML, HPC, gaming
HPC, AMD ecosystem
Intel ecosystem
Apple ecosystem
Reality Check: NVIDIA CUDA adalah standar de facto untuk ML. ROCm growing cepat (MI300X kompetitif) tapi masih ada gap. oneAPI dan Metal terbatas di ekosistem masing-masing.