Cookbook Blueprint
Code Assistant
Train a specialized coding companion on your proprietary internal libraries, API client packages, and framework conventions.
Training Data Best Practices
- Include PEP 484 Type Hints: Models trained on typed signatures hallucinate significantly fewer type mismatches.
- Provide Complete Error Handling: Include
try / exceptblocks and defensive edge-case checks in target completions. - Incorporate Sphinx/Google Docstrings: Ensures the model outputs comprehensive documentation for all generated functions.
Sample Training Data Format (JSONL)
cat > data/raw/code_pairs.jsonl << 'EOF'
{"messages": [{"role": "user", "content": "Write a Python function to reverse a string with type annotations"}, {"role": "assistant", "content": "def reverse_string(s: str) -> str:\n \"\"\"Reverse an input string safely.\"\"\"\n return s[::-1]\n\n# Example usage:\nassert reverse_string('hello') == 'olleh'"}]}
{"messages": [{"role": "user", "content": "Explain this code: for i in range(10): print(i)"}, {"role": "assistant", "content": "This loop iterates over integers from 0 to 9 (exclusive of 10) and prints each integer to stdout using standard Python range semantics."}]}
EOF
Evaluation Suite (eval_code.yaml)
name: "code-assistant-eval"
tests:
- id: "test_factorial"
prompt: "Write a recursive Python function to compute the factorial of n."
expected_keywords:
- "def factorial"
- "if n <= 1:"
- "return"
forbidden_terms:
- "goto"
- "as an AI model"
Execution & Deployment
# 1. Compile dataset moro data build --source ./data/raw/code_pairs.jsonl # 2. Train on local GPU moro train --model Qwen/Qwen2.5-Coder-1.5B-Instruct # 3. Export to GGUF and deploy to Ollama moro release export --run-id best --version v1.0.0 moro release deploy --target ollama --version v1.0.0 --name code-companion