Quickstart: 5 Minutes to Your First Local Model
This step-by-step tutorial takes you from zero to a deployed, private local language model in 5 minutes.
Step 1: Install MoroAI (30 seconds)
Install the MoroAI core package with training extensions:
# Install with training support pip install "moroai[train]"
Verify that your environment and GPU are ready:
moro doctor
Step 2: Initialize Your Project (30 seconds)
moro init my-first-model cd my-first-model
This creates a reproducible workspace:
my-first-model/ ├── moro.yaml # Project configuration & recipe ├── data/ │ └── raw/ # Raw training data inputs ├── runs/ # Training checkpoints & telemetry ├── releases/ # Exported GGUF models & SBOMs └── eval/ # Evaluation suites
Step 3: Add Your Data (1 minute)
Create a sample customer support dataset:
cat > data/raw/support_data.jsonl << 'EOF'
{"messages": [{"role": "user", "content": "How do I reset my password?"}, {"role": "assistant", "content": "Go to Settings > Security > Reset Password."}]}
{"messages": [{"role": "user", "content": "What is your refund policy?"}, {"role": "assistant", "content": "We offer a 30-day money-back guarantee on all purchases."}]}
{"messages": [{"role": "user", "content": "How do I contact support?"}, {"role": "assistant", "content": "You can reach our support team at support@example.com."}]}
EOF
Ensure moro.yaml references your raw data:
dataset: source: ./data/raw/support_data.jsonl format: jsonl
Step 4: Build the Dataset with MI Guard (30 seconds)
moro data build
✓ Loaded 3 samples ✓ After deduplication: 3 samples ✓ MI Guard: 1 rare domain landmark protected ✓ Split: train=2, val=1, eval=0 ✓ Dataset compiled to data/compiled
Step 5: Train Your Model (2 minutes)
moro train
MoroAI detects your hardware, audits VRAM memory safety, and runs local fine-tuning:
✓ Hardware detected: NVIDIA GeForce RTX 4070 (12GB) ✓ Recipe generated: lora_r=16, lr=2e-4, batch=2, accum=8 ✓ Training started: run_1234567890 ✓ Step 50/100 | Loss: 1.42 | VRAM: 6.2GB ✓ Training completed in 45s (Zero OOM)
Step 6: Deploy to Ollama (30 seconds)
moro release export --run-id run_1234567890 --version v1.0.0 moro release deploy --target ollama --version v1.0.0
Step 7: Test Your Private Model (30 seconds)
# Test locally via Ollama CLI:
ollama run my-first-model:v1.0.0 "How do I reset my password?"
# Or via the OpenAI-compatible local REST API:
curl http://localhost:11434/api/generate -d '{
"model": "my-first-model:v1.0.0",
"prompt": "How do I reset my password?",
"stream": false
}'
What Just Happened?
Raw data was parsed, deduplicated, and scored with information-theoretic token density.
An optimal LoRA recipe was synthesized to guarantee zero out-of-memory crashes on your GPU.
Adapter weights were quantized to GGUF format and registered directly into the local Ollama runtime.
Zero tokens, weights, or logs left your workstation. Complete data sovereignty achieved.