Cookbook Blueprint

Code Assistant

Train a specialized coding companion on your proprietary internal libraries, API client packages, and framework conventions.

Training Data Best Practices

  1. Include PEP 484 Type Hints: Models trained on typed signatures hallucinate significantly fewer type mismatches.
  2. Provide Complete Error Handling: Include try / except blocks and defensive edge-case checks in target completions.
  3. Incorporate Sphinx/Google Docstrings: Ensures the model outputs comprehensive documentation for all generated functions.

Sample Training Data Format (JSONL)

cat > data/raw/code_pairs.jsonl << 'EOF'
{"messages": [{"role": "user", "content": "Write a Python function to reverse a string with type annotations"}, {"role": "assistant", "content": "def reverse_string(s: str) -> str:\n    \"\"\"Reverse an input string safely.\"\"\"\n    return s[::-1]\n\n# Example usage:\nassert reverse_string('hello') == 'olleh'"}]}
{"messages": [{"role": "user", "content": "Explain this code: for i in range(10): print(i)"}, {"role": "assistant", "content": "This loop iterates over integers from 0 to 9 (exclusive of 10) and prints each integer to stdout using standard Python range semantics."}]}
EOF
  

Evaluation Suite (eval_code.yaml)

name: "code-assistant-eval"
tests:
  - id: "test_factorial"
    prompt: "Write a recursive Python function to compute the factorial of n."
    expected_keywords:
      - "def factorial"
      - "if n <= 1:"
      - "return"
    forbidden_terms:
      - "goto"
      - "as an AI model"
  

Execution & Deployment

# 1. Compile dataset
moro data build --source ./data/raw/code_pairs.jsonl

# 2. Train on local GPU
moro train --model Qwen/Qwen2.5-Coder-1.5B-Instruct

# 3. Export to GGUF and deploy to Ollama
moro release export --run-id best --version v1.0.0
moro release deploy --target ollama --version v1.0.0 --name code-companion