Subsystem 02

VRAM Recipe Predictor

Eliminating trial-and-error fine-tuning by simulating memory consumption down to the megabyte before touching GPU memory.

The VRAM Memory Formula

When fine-tuning on consumer hardware, memory exhaustion is non-linear. Peak GPU VRAM is modeled prior to training execution across five critical tensor allocations:

Analytical VRAM Memory Model Peak Memory Equation
VRAM_Peak = M_base + M_lora + M_optimizer + M_activations + M_cuda_context

Where:
  M_base        = (N_params * B_quant) / 8           # Base weights in NF4 / FP16
  M_lora        = 2 * rank * d_model * N_targets * 4 # Trainable LoRA adapter matrices (FP32)
  M_optimizer   = M_lora * 2                         # 8-bit AdamW first and second moments
  M_activations = b * s * L * (34 * d_model + 5 * s * h) # Backprop activation tensors
  M_cuda_ctx    = ~850MB to 1.2GB                     # CUDA driver context & kernels
    

Hardware Tier Presets & Validation

The engine maintains validated memory envelopes across standard consumer and workstation GPUs:

Hardware Profile Available Memory Recommended Model LoRA Rank (r) Context Length
RTX 3060 / 4060 12 GB VRAM Qwen2.5-1.5B / Llama-3.2-1B r=16, alpha=32 2,048 tokens
RTX 4070 / 4070 Ti 12–16 GB VRAM Qwen2.5-3B / Llama-3.2-3B r=32, alpha=64 4,096 tokens
RTX 3090 / 4090 24 GB VRAM Qwen2.5-7B / Llama-3.1-8B r=64, alpha=128 4,096 tokens
Apple Silicon (M2/M3/M4) 16–64 GB Unified Qwen2.5-7B / Mistral-7B r=16, alpha=32 2,048 tokens
Cloud A100 / H100 80 GB VRAM Qwen2.5-14B / 32B r=64, alpha=128 8,192 tokens

Generating a Reproducible Recipe

Run the recipe engine to automatically probe your active GPU driver, detect memory bandwidth, and output a guaranteed zero-OOM moro.yaml configuration:

# Automatically audit GPU and generate optimal training configuration
moro recipe generate --model Qwen/Qwen2.5-1.5B --hardware detect --output ./moro.yaml
  

Generated Recipe File Anatomy (moro.yaml)

name: "customer-support-qwen"
version: "0.1.0"
model:
  base: "Qwen/Qwen2.5-1.5B"
  quantization: "nf4"        # 4-bit NormalFloat quantization
  lora_rank: 16
  lora_alpha: 32
  lora_dropout: 0.05
  target_modules: ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]

training:
  micro_batch_size: 2
  gradient_accumulation_steps: 8
  learning_rate: 2.0e-4
  max_seq_length: 2048
  optimizer: "adamw_8bit"
  gradient_checkpointing: true
  auto_heal: true            # Enables 5-level autonomous OOM recovery