Local Serving
Deploying to Local Ollama
Export your fine-tuned LoRA adapters directly into Ollama with GGUF quantization in a single command.
Prerequisites
- Ollama installed and running locally (
ollama --version) - A completed MoroAI training run with a certified checkpoint
1-Command Deployment
moro release deploy --target ollama --quantize q4_k_m --model-name my-assistant:v1
MoroAI executes the following steps behind the scenes:
- Merges the low-rank adapter weights into the base model tensors.
- Executes
llama.cppquantization intoQ4_K_MGGUF format. - Generates a tailored
Modelfilewith system prompt templates and stop tokens. - Calls the Ollama API to create and register the model tag locally.
Testing Your Model
ollama run my-assistant:v1 "Hello! Summarize our return policy."