Lade Stack
Architecture

System Architecture

A deep technical breakdown of how Lade Stack is designed. From model routing to GPU serving, every layer is built for performance, reliability, and cost efficiency.

Overview

High-Level System Design

The system is composed of four primary layers: CLI interface, routing engine, model serving, and GPU infrastructure.

CLI Layer

  • Command Parser
  • Context Loader
  • Config Manager
  • Output Formatter

Routing Layer

  • Task Classifier
  • Model Selector
  • Cost Optimizer
  • Fallback Handler

Serving Layer

  • vLLM Runtime
  • Batch Scheduler
  • KV Cache Manager
  • Token Streamer

Infrastructure

  • GPU Pool Manager
  • Auto Scaler
  • Health Monitor
  • Model Registry
CLIRouterServingGPU

Model Routing Layer

The routing engine classifies incoming tasks and selects the optimal model based on task type, context length, latency requirements, and cost constraints. Routing decisions are transparent and logged for debugging.

routing-config.yaml
routing:
  strategy: "task-optimized"
  rules:
    - task: code_generation
      primary: qwen-coder-32b
      fallback: deepseek-coder-v2
      max_latency_ms: 2000

    - task: debugging
      primary: deepseek-r1
      fallback: qwen-coder-32b
      max_latency_ms: 5000

    - task: long_context
      primary: kimi-128k
      context_threshold: 32000
      max_latency_ms: 10000

    - task: multilingual
      primary: glm-4
      fallback: qwen-coder-32b

GPU Serving Layer

Models are served through vLLM or TensorRT-LLM for maximum throughput. Continuous batching, PagedAttention, and speculative decoding minimize latency and maximize GPU utilization.

serving-config.yaml
serving:
  engine: vllm
  config:
    tensor_parallel_size: 2
    max_num_batched_tokens: 32768
    gpu_memory_utilization: 0.90
    enable_chunked_prefill: true
    max_model_len: 131072

  models:
    - name: qwen-coder-32b
      quantization: awq-4bit
      gpu_count: 2
      max_concurrent: 16

    - name: deepseek-r1
      quantization: gptq-8bit
      gpu_count: 4
      max_concurrent: 8

Auto Scaling Strategy

The auto scaler monitors request queue depth, GPU utilization, and latency percentiles to make scaling decisions. Scale-up is aggressive to minimize wait times. Scale-down uses a cooldown period to prevent thrashing.

autoscaler-policy.yaml
autoscaling:
  min_replicas: 1
  max_replicas: 8
  metrics:
    - type: queue_depth
      target: 10
      scale_up_threshold: 20
      scale_down_threshold: 3

    - type: gpu_utilization
      target: 0.75
      scale_up_threshold: 0.85
      scale_down_threshold: 0.40

    - type: p99_latency_ms
      scale_up_threshold: 5000

  cooldown:
    scale_up: 60s
    scale_down: 300s

Quantization Strategy

Models are quantized to reduce memory footprint and improve throughput without significant quality loss. AWQ 4-bit is the default for coding models. 8-bit GPTQ is used for reasoning models where precision matters more.

AWQ 4-bit~50% reduction

Code generation models — < 1% degradation on HumanEval

GPTQ 8-bit~25% reduction

Reasoning models — < 0.5% degradation on MMLU

FP16Baseline

Critical accuracy tasks — Full precision

Multi-Model Pool Design

The model pool maintains warm instances of frequently used models and cold-starts less common ones on demand. Models share GPU memory through intelligent scheduling and preemption policies.

model-pool.yaml
pool:
  warm_models:
    - qwen-coder-32b    # Always loaded
    - deepseek-r1       # Always loaded

  cold_models:
    - glm-4             # Load on demand
    - kimi-128k         # Load on demand

  scheduling:
    strategy: priority_preemptive
    preemption_policy: recompute
    max_loading_concurrent: 2

  memory:
    shared_gpu_memory: true
    swap_space: 32GB
    cache_dtype: auto

Latency Optimization

Multiple techniques are employed to minimize time-to-first-token and overall response latency across the inference pipeline.

Continuous Batching

New requests are added to running batches without waiting for batch completion

PagedAttention

KV cache is managed in pages, eliminating memory fragmentation and enabling larger batch sizes

Speculative Decoding

Small draft model generates candidates that the main model verifies in parallel

Prefix Caching

Common system prompts and workspace context are cached across requests

Chunked Prefill

Long prompts are processed in chunks to interleave with decode steps from other requests

Cost Optimization

Infrastructure costs are optimized through a combination of spot instances, right-sizing, quantization, and intelligent scheduling.

Spot Instance Integration

Non-critical workloads run on spot/preemptible instances with automatic failover to on-demand

Right-Sizing

GPU allocation matches model requirements — no paying for A100s when an L4 will do

Request Batching

Concurrent requests share GPU compute through continuous batching, maximizing utilization

Idle Shutdown

Unused model instances are unloaded after configurable idle periods to free GPU memory

Roadmap

Future Architecture

Planned architectural enhancements that will expand capability and efficiency.

Fine-Tuning Pipeline

On-infrastructure fine-tuning with LoRA adapters. Train domain-specific models on your own data without leaving your environment.

Model Distillation

Distill large models into smaller, faster variants optimized for your specific use cases. Reduce latency and cost while maintaining quality.

Hybrid Fallback

Optional fallback to external API providers when local infrastructure is at capacity. Configurable cost and privacy thresholds.