πŸ€– LLM AI Systems Hierarchy β€” The Complete Production Stack

Vault QuCrypto punya 55 catatan tentang AI/LLM/RAG/Agent/MCP β€” tapi tidak ada satu dokumen pun yang menunjukkan layer architecture lengkapnya. Catatan ini memetakan stack LLM modern ke dalam 7 lapisan fungsional (dari hardware sampai product), dengan timeline evolusi 2020-2026, trade-off matriks per layer, dan diagram koneksi ke setiap catatan terkait di vault.


Daftar Isi

  1. 1. Premise β€” Mengapa Hierarchy Ini Penting
  2. 2. The Seven-Layer Model
  3. 3. Layer 0 β€” Hardware Foundation
  4. 4. Layer 1 β€” Model Foundation (Pre-training)
  5. 5. Layer 2 β€” Model Adaptation (Fine-tuning)
  6. 6. Layer 3 β€” Inference Infrastructure
  7. 7. Layer 4 β€” Context Engineering (RAG, Prompts, Memory)
  8. 8. Layer 5 β€” Agentic Layer (Tools, MCP, Multi-Agent)
  9. 9. Layer 6 β€” Evaluation & Observability
  10. 10. Layer 7 β€” Product/UX Surface
  11. 11. Timeline 2020-2026 β€” Bagaimana Kita Sampai di Sini
  12. 12. Trade-off Matrix per Layer
  13. 13. Cross-Reference ke Vault
  14. References

1. Premise β€” Mengapa Hierarchy Ini Penting

LLM bukan satu produk β€” ia adalah tumpukan 7 lapisan yang masing-masing punya disiplin sendiri:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Layer 7: Product/UX Surface                      β”‚ ← Yang dilihat user
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 6: Evaluation & Observability              β”‚ ← Bagaimana kita tau ini jalan?
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 5: Agentic Layer                           β”‚ ← Tools, MCP, planning
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 4: Context Engineering                     β”‚ ← RAG, prompts, memory
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 3: Inference Infrastructure               β”‚ ← Serving, batching, KV-cache
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 2: Model Adaptation                       β”‚ ← Fine-tuning, LoRA, adapters
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 1: Model Foundation                       β”‚ ← Pre-training, architecture
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer 0: Hardware Foundation                    β”‚ ← GPU, TPU, memory, interconnect
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Mengapa 7 layer? Karena setiap layer punya trade-off independen: hardware menentukan model mana yang bisa dilatih, model menentukan inference pattern, fine-tuning menentukan capability, dst. Salah optimasi di salah satu layer = sia-sia di layer lain.

Mengapa vault butuh hierarchy ini? Karena catatan AI_Systems di vault (hierarchy-kernel-bypass-networking, llm-finetuning-toolchain, agentic-ai-mcp-architecture-deepdive, dll.) membahas layer-layer individual tapi tidak ada yang memetakan semua layer dalam satu diagram. Catatan ini jadi peta.


2. The Seven-Layer Model

2.1 Definisi Setiap Layer

LayerFungsiOwnerFailure Mode Tipikal
0HardwareNVIDIA, AMD, TPUs, custom ASICGPU out of memory, interconnect bottleneck
1Model foundationOpenAI, Anthropic, Meta, DeepSeekLoss tidak konvergen, hallucination
2AdaptationDownstream teamsCatastrophic forgetting, overfit
3InferenceOps/DevOpsTTFT p99 > 3s, cost per token meledak
4ContextAI/ML engineersRetrieval recall <80%, prompt injection
5AgenticApplication devsInfinite loop, tool hallucination
6EvaluationQA + researchRegression drift silent, rogue eval
7Product/UXDesigners + PMUser trust eroded, churn

2.2 Layer Dependency

[Layer N] ←————————————————— depends on β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β†’ [Layer N-1]

Layer 7 butuh             β†’ Layer 6 (metrics) + Layer 5 (agent flow)
Layer 6 butuh             β†’ Layer 3 (traces) + Layer 5 (decision logs)
Layer 5 butuh             β†’ Layer 4 (context) + Layer 3 (latency budget)
Layer 4 butuh             β†’ Layer 3 (inference cost) + Layer 2 (capability)
Layer 3 butuh             β†’ Layer 2 (model) + Layer 1 (weights) + Layer 0 (GPU)
Layer 2 butuh             β†’ Layer 1 (base model)
Layer 1 butuh             β†’ Layer 0 (compute cluster)

Prinsip: Optimasi di Layer N tidak bisa mengkompensasi kelemahan di Layer N-1. Contoh: prompt yang indah di Layer 4 tidak akan menyelamatkan model yang cacat di Layer 1.


3. Layer 0 β€” Hardware Foundation

GPU/TPU + memory hierarchy + interconnect yang menentukan throughput.

3.1 Komponen Kritis

SubsystemSpesifikasiVendorThroughput
ComputeH100, H200, B200 (NVIDIA) β€” MI300X (AMD) β€” TPU v5p (Google)NVIDIA dominates 88%1-4 PFLOPs FP8
MemoryHBM3, HBM3e (up to 192 GB/stack)SK Hynix, Samsung, Micron3-6 TB/s
InterconnectNVLink (900 GB/s), InfiniBand (400 Gbps), PCIe Gen5NVIDIA, Mellanox600-900 GB/s per node
StorageNVMe SSD, parallel filesystem (Lustre, GPFS)Pure Storage, WEKA100+ GB/s read
NetworkingRoCE, RoCEv2, custom EthernetArista, Mellanox800 Gbps rollout

3.2 Compute Hierarchy

Single GPU            = H100 SXM
Node (8 GPU NVLink)   = 1.6 PFLOPs FP8 full mesh
Pod (32 nodes)        = 1024 GPU, NVLink + IB
Cluster (>1000 GPU)   = Training run skala GPT-4
Hyperscale (>100K GPU)= Frontier training (MosaicML, xAI Colossus)

Koneksi ke Vault:

3.3 Kenapa Layer 0 Penting

Setiap optimasi Layer 1 (pre-training) dibatasi oleh Layer 0. Contoh: arsitektur Mixture-of-Experts (MoE) hemat komputasi activation, tapi memperburuk memory pressure di inference. Cluster kecil β†’ batch kecil β†’ throughput rendah. Ini sebabnya vendor kecil seperti Mistral fokus di model dense < 70B yang muat di single node.


4. Layer 1 β€” Model Foundation (Pre-training)

Arsitektur model + tokenizer + pre-training objective + dataset curation.

4.1 Arsitektur yang Dominan

ArsitekturTahunInovasiUse Case
GPT-1 (decoder-only)2018Transformer decoder, causal maskingGenerative
BERT (encoder-only)2018Bidirectional MLMClassification, embedding
T5 (encoder-decoder)2019Unified text-to-textTranslation, summarization
GPT-3 (scaling)2020In-context learning emerges at 13B+Few-shot generalist
Switch Transformer2021Mixture of Experts (MoE)Sparse compute
Chinchilla2022Compute-optimal scaling law70B trained on 1.4T tokens
Llama 2/32023-24Open weights, RLHFChat, instruction following
DeepSeek-R12025RL-first reasoning modelMath, code, reasoning
Claude 3.7+2025Long context (200K-1M tokens)Document analysis

4.2 Pre-training Objective

Next-token prediction (GPT-style):

Justru objective yang sangat sederhana ini, dalam arsitektur yang tepat dengan cukup compute, menghasilkan emergent capabilities (in-context learning, chain-of-thought) yang tidak diprogram secara eksplisit.

4.3 Dataset Curation

Data SourceRasio TipikalKarakteristik
Web crawl (Common Crawl)60-80%Sangat besar tapi noisy
Books (Books3, BooksCorpus)5-15%High-quality narrative
Code (GitHub, Stack)5-15%Struktural, dapat diuji
Wikipedia2-5%Factual, multilingual
Scientific papers1-3%Domain-specific
Synthetic (GPT-4 generated)5-10%Curated instruction data

Filter pipeline: Quality β†’ Deduplication β†’ Toxicity β†’ PII removal β†’ Length filtering. Total data yang dipakai = ~10% dari raw crawl.

Koneksi ke Vault:


5. Layer 2 β€” Model Adaptation (Fine-tuning)

Mengubah perilaku/capability model untuk kasus spesifik dengan biaya komputasi yang jauh lebih rendah dari pre-training.

5.1 Teknik Fine-tuning

TeknikTahunParameter UpdateCompute vs Full FTMemori GPU
Full fine-tuning2017100% Semua parameter1Γ—100%
LoRA2021~0.1% (rank decomposition)~3Γ— lebih murah~10%
QLoRA2023LoRA + 4-bit quantization~10Γ— lebih murah~3%
Adapter (Houlsby)2019~5% (small modules injected)~5Γ— lebih murah~15%
Prompt tuning20210% (soft prompt trained)Negligible<1%
Prefix tuning20210% + KV prefix learnedNegligible<1%
DoRA2024LoRA enhanced (magnitude + direction)~3.5Γ—~12%
Full FT (MoE expert tuning)2025~10%~2Γ—~25%

5.2 Algoritma Fine-tuning Modern

RLHF (Reinforcement Learning from Human Feedback):

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Pre-trainedβ”‚ β†’  β”‚ SFT (Supervisedβ”‚ β†’ β”‚ Reward Modelβ”‚ β†’ β”‚ PPO          β”‚
β”‚ Model      β”‚    β”‚ Fine-tuning)   β”‚    β”‚ (Human-rankedβ”‚    β”‚ (RL)         β”‚
β”‚            β”‚    β”‚                β”‚    β”‚ preferences)β”‚    β”‚              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

DPO (Direct Preference Optimization) β€” 2024, lebih sederhana dari PPO, tanpa reward model terpisah:

Dimana = preferred, = rejected, = policy yang dilatih.

5.3 Dataset untuk Fine-tuning

JenisFormatContoh
SFT(instruction, response) pairsAlpaca, Dolly, OASST
Preference(prompt, chosen, rejected)HH-RLHF, UltraFeedback
Reasoning(problem, step-by-step)OpenMath, GSM8K-CoT
Tool use(query, tool_call, result)ToolBench, Hermes

Koneksi ke Vault:


6. Layer 3 β€” Inference Infrastructure

Bagaimana model yang sudah dilatih disajikan ke pengguna β€” dengan latency, throughput, dan biaya yang optimal.

6.1 Tantangan Inference LLM

Inference LLM punya profile yang unik:

Pre-training  : Latency不重要 (perlu waktu)
Fine-tuning   : Latency不重要 (perlu waktu)
Inference     : Latency P99 < 2s MUTLAK
              : Throughput tinggi (1000-10000 req/s)
              : Biaya per token minimal

Perbedaan fundamental dengan inference tradisional:

  1. Memory-bound, bukan compute-bound β€” kv-cache memenuhi seluruh GPU memory
  2. Sequence length variable β€” TTFT dan TPOT berbeda
  3. Statelessness terbatas β€” context window besar berarti stateful

6.2 Teknik Serving Modern

TeknikFungsiThroughput ImpactLatency Impact
Continuous batchingDynamic gap-filling5-20Γ—-50% TTFT
PagedAttention (vLLM)Memory paging untuk kv-cache10-24Γ—-30% TPOT
Speculative decodingDraft model + verification2-3Γ—-60% TPOT
Prefix cachingCache common system prompts10Γ— untuk RAG-90% TTFT
Quantization (FP8/INT4)Reduce memory bandwidth2-4Γ—-40% TTFT
Multi-LoRA servingServe many adapters in one GPU4Γ—+20% TTFT
Disaggregated inferenceSeparate prefill vs decode2-3Γ—-50% P99
Tensor parallelismShard besar ke banyak GPU--70% latency per token
Pipeline parallelismMulti-node sharding-Communication bound

6.3 Production Stack

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Application Layer (prompt β†’ stream response)    β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Scheduler (rate limiting, queue, SLO)           β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Inference Engine (vLLM / TGI / TensorRT-LLM)    β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Model Optimizer (compiler, quantizer, pruner)   β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ GPU Runtime (CUDA / ROCm, MIG/MPS)              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Production Choice Matrix (2025):

EngineThroughput (tok/s/GPU)HardwareBahasa
vLLM8000-15000NVIDIA, AMDPython/CUDA
TGI (HuggingFace)6000-12000NVIDIARust/Python
TensorRT-LLM12000-25000NVIDIAPython/C++/Triton
SGLang15000-30000NVIDIAPython/Rust
llama.cpp (CPU)50-500CPU-onlyC++
MLX (Apple Silicon)800-3000M-seriesPython/C++

Koneksi ke Vault:


7. Layer 4 β€” Context Engineering (RAG, Prompts, Memory)

Memberikan model informasi yang relevan untuk menjawab query β€” melampaui apa yang ada di parameter.

7.1 The Context Engineering Stack

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ User Query + History + State                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Query Transformation                               β”‚
β”‚ (rewrite, decompose, HyDE, step-back)              β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Retrieval (BM25 + dense + hybrid + reranker)       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Context Assembly                                   β”‚
β”‚ (windowing, hierarchy, lost-in-the-middle)         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Prompt Composition                                 β”‚
β”‚ (system + retrieved + user)                        β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Response Generation                                β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

7.2 RAG Stages

StageTeknikTrade-off
IngestionChunker, embedder, metadata extractorChunk size vs recall
Query transformRewrite, query expansion, HyDERecall vs latency
RetrievalBM25 / dense / hybrid / multi-vectorRecall vs precision
RerankingCross-encoder, LLM-based, ColBERTLatency vs MRR
GenerationLLM final answer with contextContext window vs accuracy

7.3 Chunking Strategies

StrategyChunk SizeOverlapRecallSpeed
Fixed-size256-10240-100SedangCepat
Recursive splitter256-10240-100BagusCepat
SemanticVariable0TerbaikLambat
Document-awareVariable0TerbaikSedang

7.4 Advanced RAG Patterns

  • Self-RAG: Model decide sendiri apakah perlu retrieval
  • Corrective RAG (CRAG): Hasil retrieval di-grade, lalu re-retrieve jika buruk
  • Agentic RAG: Multi-hop retrieval dengan planning
  • GraphRAG: Graph-based retrieval untuk relationship queries
  • Multi-modal RAG: Text + image + audio + video dalam satu sistem

Koneksi ke Vault:


8. Layer 5 β€” Agentic Layer (Tools, MCP, Multi-Agent)

Memberi model kemampuan untuk bertindak di dunia dengan memanggil tools, multi-step planning, dan kolaborasi multi-agent.

8.1 Tool Use Evolution

EvolusiPatternContoh
1.0ReAct (Reason + Act)LangChain agents 2022
2.0Tool calling (native)OpenAI function calling 2023
3.0Structured outputJSON Mode, Pydantic, Zod
4.0MCP (Model Context Protocol)Anthropic/KhΓΊc 2024
5.0Multi-agent collaborationAutoGen, CrewAI 2024
6.0Autonomous codingClaude Code, Codex, Devin
7.0Persistent memory + SkillsHermes agents 2025-26

8.2 MCP Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  protocol   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  calls   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ LLM Host     β”‚ ←─JSON-RPC─→│ MCP Server   β”‚ ──────→  β”‚ External   β”‚
β”‚ (Claude Code)β”‚              β”‚ (Filesystem) β”‚          β”‚ Service    β”‚
β”‚              β”‚              β”‚              β”‚          β”‚            β”‚
β”‚ Tools: list  β”‚              β”‚ Tools: read  β”‚          β”‚            β”‚
β”‚ Resources:   β”‚              β”‚ Resources:   β”‚          β”‚            β”‚
β”‚  list        β”‚              β”‚  list        β”‚          β”‚            β”‚
β”‚ Prompts:     β”‚              β”‚ Prompts:     β”‚          β”‚            β”‚
β”‚  list        β”‚              β”‚  templates   β”‚          β”‚            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Three primitives:

  1. Tools β€” model-controlled actions
  2. Resources β€” application-controlled context
  3. Prompts β€” user-controlled templates

8.3 Multi-Agent Patterns

PatternAgenUse Case
Supervisor-Worker1 supervisor + N workersTask decomposition
Peer-to-PeerNεΉ³η­‰ηš„ agenDebate, voting
HierarchicalTree of agentsComplex workflows
BlackboardShared memorySpecialist collaboration
SwarmLightweight coordinationEphemeral tasks
CrewRole-based teamSoftware dev (CrewAI)

Koneksi ke Vault:


9. Layer 6 β€” Evaluation & Observability

Bagaimana kita tahu kalau sistem AI kita sesuai spec dan tetap sesuai spec seiring waktu.

9.1 Three-Layer Evaluation

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Layer C: User Feedback              β”‚ ← Real users (thumbs up/down)
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer B: Online Metrics             β”‚ ← Production traces (TTFT, tokens, refusal)
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Layer A: Offline Evals              β”‚ ← Curated datasets (accuracy, bias, safety)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

9.2 Offline Evaluation Categories

KategoriMetrikContoh
CapabilityMMLU, GSM8K, HumanEvalGeneral skill
DomainMedQA, LegalBenchSpecialist
SafetyHarmBench, AdvBenchAdversarial inputs
BiasBBQ, StereoSetFairness
ReasoningARC, BBHMulti-step
Instruction followingIFEval, Multi-IFFormat adherence
RetrievalnDCG@10, Recall@10Untuk RAG
AgentTool accuracy, task successTool use

9.3 Observability Stack

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Distributed Tracing (OpenTelemetry) β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Token usage & cost analytics         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ LLM-as-judge for automated eval     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Drift detection (embedding + output)β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Cost guardrails + rate limiting     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Koneksi ke Vault:


10. Layer 7 β€” Product/UX Surface

Bagaimana AI disajikan ke user akhir β€” chat box, voice, autonomous action, atau bentuk lainnya.

10.1 UX Patterns

PatternLatency BudgetUse CaseContoh
Streaming chatTTFT <500msConversationalChatGPT, Claude
Voice (live)TTFT <200msSpoken conversationElevenLabs, Gemini Live
Batch asyncTidak real-timeLong document analysisNotebookLM
Agent asyncMinutes-hoursAutonomous codingDevin, Claude Code
Inline assist<100msCode completionCopilot, Cursor
Predictive<50msAuto-completeSmart reply
Multi-modal<1sRealtime vision/audioGPT-4o realtime

10.2 Trust & Safety Surface

Setiap layer harus transparan ke user:

  • Citations (Layer 4)
  • Confidence display (Layer 6)
  • Action preview (Layer 5)
  • Reasoning trace (Layer 5 eval)
  • Cost display (Layer 3 + 6)

Koneksi ke Vault:


11. Timeline 2020-2026 β€” Bagaimana Kita Sampai di Sini

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Year β”‚ Major Milestone                                    β”‚
β”œβ”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 2020 β”‚ GPT-3 (175B) β€” in-context learning emerges       β”‚
β”‚ 2021 β”‚ Codex, Copilot β€” coding becomes mainstream       β”‚
β”‚ 2022 β”‚ ChatGPT (RLHF), Stable Diffusion, Whisper        β”‚
β”‚      β”‚ β†’ Layer 5 Architecture stabilizes                 β”‚
β”‚ 2023 β”‚ GPT-4, Claude 2, Llama 2 β€” multimodal, long ctx  β”‚
β”‚      β”‚ β†’ Layer 1-3 production-grade                      β”‚
β”‚      β”‚ β†’ RAG becomes default pattern (Layer 4)          β”‚
β”‚      β”‚ β†’ Function calling native (Layer 5 v2)            β”‚
β”‚ 2024 β”‚ Claude 3.5 Sonnet, Llama 3, DeepSeek-V3          β”‚
β”‚      β”‚ β†’ MCP protocol standardized                      β”‚
β”‚      β”‚ β†’ Multi-agent frameworks mature                   β”‚
β”‚      β”‚ β†’ QLoRA democratizes Layer 2                     β”‚
β”‚      β”‚ β†’ Continuous batching di vLLM (Layer 3)           β”‚
β”‚ 2025 β”‚ Claude 4, GPT-5, DeepSeek-R1 β€” reasoning models  β”‚
β”‚      β”‚ β†’ Test-time compute scaling (chain-of-thought)    β”‚
β”‚      β”‚ β†’ Long context >1M tokens                        β”‚
β”‚      β”‚ β†’ Agentic coding mainstream                       β”‚
β”‚      β”‚ β†’ Reasoning tokens visible (CoT)                  β”‚
β”‚ 2026 β”‚ Multi-modal real-time, persistent memory         β”‚
β”‚      β”‚ β†’ Skills marketplace emerging                     β”‚
β”‚      β”‚ β†’ Self-improving agents (RL on production)        β”‚
β”‚      β”‚ β†’ MCP 2.0 β€” federated agent networks              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

12. Trade-off Matrix per Layer

Setiap layer punya trade-off Cost vs Capability vs Latency:

LayerCost DriverCapability LeverLatency Lever
0GPU hours Γ— $/hourModel sizeCluster interconnect
1Training tokens Γ— energyArchitecture innovationN/A (offline)
2Adapter size Γ— datasetDataset qualityTraining compute
3GPU-seconds Γ— batch sizeModel qualityBatching strategy
4Embedding + LLM callReranker qualityIndex structure
5Tool calls Γ— costPlanning depthMax iterations
6Eval sampling rateEval dataset diversityCaching strategy
7User satisfactionSurface clarityStreaming

Kontradiksi utama:

  • Layer 3 (inference cost) β™₯ tekanan dari Layer 7 (latency) tapi βœ— konflik dengan Layer 4 (recall)
  • Layer 5 (agent capability) β™₯ keinginan Layer 7 (autonomy) tapi βœ— konflik dengan Layer 6 (predictability)
  • Layer 2 (capability) β™₯ keinginan Layer 4 (retrieval quality) tapi βœ— konflik dengan Layer 3 (latency)

13. Cross-Reference ke Vault

Vault QuCrypto sudah punya catatan di masing-masing layer. Daftar berikut memetakan vault notes β†’ layer:


References

  1. Vaswani et al. β€œAttention Is All You Need.” NeurIPS 2017.
  2. Brown et al. β€œLanguage Models are Few-Shot Learners.” arXiv:2005.14165 (2020).
  3. Hu et al. β€œLoRA: Low-Rank Adaptation of Large Language Models.” ICLR 2022.
  4. Dettmers et al. β€œQLoRA: Efficient Finetuning of Quantized LLMs.” NeurIPS 2023.
  5. Ouyang et al. β€œTraining Language Models to Follow Instructions with Human Feedback.” (2022).
  6. Rafailov et al. β€œDirect Preference Optimization.” NeurIPS 2023.
  7. Kwon et al. β€œEfficient Memory Management for Large Language Model Serving with PagedAttention.” SOSP 2023.
  8. Anthropic. β€œModel Context Protocol Specification.” (2024). https://modelcontextprotocol.io/
  9. Park et al. β€œGenerative Agents: Interactive Simulacra of Human Behavior.” UIST 2023.
  10. Kapoor et al. β€œAI Evaluations: A Taxonomy of What to Evaluate.” Stanford HAI (2024).
  11. Anthropic. β€œClaude 3.7 System Card.” (2025).
  12. DeepSeek Team. β€œDeepSeek-R1: Incentivizing Reasoning Capability.” (2025).
  13. S. Borgeaud et al. β€œImproving Language Models by Retrieving from Trillions of Tokens.” ICML 2022.
  14. J. Wei et al. β€œChain-of-Thought Prompting Elicits Reasoning in Large Language Models.” NeurIPS 2022.
  15. Patrick Lewis et al. β€œRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” NeurIPS 2020.