System Architecture
A deep technical breakdown of how Lade Stack is designed. From model routing to GPU serving, every layer is built for performance, reliability, and cost efficiency.
High-Level System Design
The system is composed of four primary layers: CLI interface, routing engine, model serving, and GPU infrastructure.
CLI Layer
- Command Parser
- Context Loader
- Config Manager
- Output Formatter
Routing Layer
- Task Classifier
- Model Selector
- Cost Optimizer
- Fallback Handler
Serving Layer
- vLLM Runtime
- Batch Scheduler
- KV Cache Manager
- Token Streamer
Infrastructure
- GPU Pool Manager
- Auto Scaler
- Health Monitor
- Model Registry
Model Routing Layer
The routing engine classifies incoming tasks and selects the optimal model based on task type, context length, latency requirements, and cost constraints. Routing decisions are transparent and logged for debugging.
routing:
strategy: "task-optimized"
rules:
- task: code_generation
primary: qwen-coder-32b
fallback: deepseek-coder-v2
max_latency_ms: 2000
- task: debugging
primary: deepseek-r1
fallback: qwen-coder-32b
max_latency_ms: 5000
- task: long_context
primary: kimi-128k
context_threshold: 32000
max_latency_ms: 10000
- task: multilingual
primary: glm-4
fallback: qwen-coder-32bGPU Serving Layer
Models are served through vLLM or TensorRT-LLM for maximum throughput. Continuous batching, PagedAttention, and speculative decoding minimize latency and maximize GPU utilization.
serving:
engine: vllm
config:
tensor_parallel_size: 2
max_num_batched_tokens: 32768
gpu_memory_utilization: 0.90
enable_chunked_prefill: true
max_model_len: 131072
models:
- name: qwen-coder-32b
quantization: awq-4bit
gpu_count: 2
max_concurrent: 16
- name: deepseek-r1
quantization: gptq-8bit
gpu_count: 4
max_concurrent: 8Auto Scaling Strategy
The auto scaler monitors request queue depth, GPU utilization, and latency percentiles to make scaling decisions. Scale-up is aggressive to minimize wait times. Scale-down uses a cooldown period to prevent thrashing.
autoscaling:
min_replicas: 1
max_replicas: 8
metrics:
- type: queue_depth
target: 10
scale_up_threshold: 20
scale_down_threshold: 3
- type: gpu_utilization
target: 0.75
scale_up_threshold: 0.85
scale_down_threshold: 0.40
- type: p99_latency_ms
scale_up_threshold: 5000
cooldown:
scale_up: 60s
scale_down: 300sQuantization Strategy
Models are quantized to reduce memory footprint and improve throughput without significant quality loss. AWQ 4-bit is the default for coding models. 8-bit GPTQ is used for reasoning models where precision matters more.
Code generation models — < 1% degradation on HumanEval
Reasoning models — < 0.5% degradation on MMLU
Critical accuracy tasks — Full precision
Multi-Model Pool Design
The model pool maintains warm instances of frequently used models and cold-starts less common ones on demand. Models share GPU memory through intelligent scheduling and preemption policies.
pool:
warm_models:
- qwen-coder-32b # Always loaded
- deepseek-r1 # Always loaded
cold_models:
- glm-4 # Load on demand
- kimi-128k # Load on demand
scheduling:
strategy: priority_preemptive
preemption_policy: recompute
max_loading_concurrent: 2
memory:
shared_gpu_memory: true
swap_space: 32GB
cache_dtype: autoLatency Optimization
Multiple techniques are employed to minimize time-to-first-token and overall response latency across the inference pipeline.
Continuous Batching
New requests are added to running batches without waiting for batch completion
PagedAttention
KV cache is managed in pages, eliminating memory fragmentation and enabling larger batch sizes
Speculative Decoding
Small draft model generates candidates that the main model verifies in parallel
Prefix Caching
Common system prompts and workspace context are cached across requests
Chunked Prefill
Long prompts are processed in chunks to interleave with decode steps from other requests
Cost Optimization
Infrastructure costs are optimized through a combination of spot instances, right-sizing, quantization, and intelligent scheduling.
Spot Instance Integration
Non-critical workloads run on spot/preemptible instances with automatic failover to on-demand
Right-Sizing
GPU allocation matches model requirements — no paying for A100s when an L4 will do
Request Batching
Concurrent requests share GPU compute through continuous batching, maximizing utilization
Idle Shutdown
Unused model instances are unloaded after configurable idle periods to free GPU memory
Future Architecture
Planned architectural enhancements that will expand capability and efficiency.
Fine-Tuning Pipeline
On-infrastructure fine-tuning with LoRA adapters. Train domain-specific models on your own data without leaving your environment.
Model Distillation
Distill large models into smaller, faster variants optimized for your specific use cases. Reduce latency and cost while maintaining quality.
Hybrid Fallback
Optional fallback to external API providers when local infrastructure is at capacity. Configurable cost and privacy thresholds.