Multi-Layer Evaluation Harness
Comprehensive model certification across deterministic rules, semantic judges, distribution drift detection, and adversarial stress tests.
Beyond Validation Loss: The Need for Multi-Layer Certification
In academic research, evaluation often stops at cross-entropy validation loss or standard perplexity. However, in enterprise deployment, an adapted model with a low validation loss can still:
- Violate strict JSON schema requirements, producing malformed API payloads
- Hallucinate plausible-sounding legal or medical falsehoods
- Experience catastrophic forgetting of general reasoning capabilities
- Break under noisy user input with misspellings or colloquial phrasing
MoroAI enforces four distinct evaluation gates before certifying any checkpoint for production release.
The 4-Layer Validation Gates
Validates completions against strict deterministic patterns. Checks include exact JSON Schema validation, regex constraints for mandatory IDs or formatting, and zero-tolerance scans for forbidden tokens or hallucinated URLs.
Executes reference-guided domain evaluation against ground-truth question-answer pairs. An evaluator model scores factual correctness, reasoning completeness, and tone fidelity on a calibrated 1-5 scale using structured JSON rubrics.
Calculates the Earth Mover's (Wasserstein-1) distance between output embedding representations of the fine-tuned model versus the base foundation model on standard anchor prompts. If drift exceeds $W_1 > 0.05$, the model is flagged for catastrophic forgetting.
Automatically injects synthetic perturbations into test prompts (character swaps, common keyboard typos, and homoglyphs). The harness verifies that the model's structured completion remains invariant under realistic noisy operator input.
Defining an Evaluation Suite (eval_suite.yaml)
name: "customer-support-eval"
version: "1.0.0"
gates:
deterministic:
min_pass_rate: 1.0 # 100% must pass JSON schema
semantic:
min_accuracy: 0.90 # 90% domain accuracy
drift:
max_wasserstein: 0.05 # No severe representation collapse
adversarial:
min_stability: 0.85 # Robust against typos
tests:
- id: "test_refund_json"
prompt: "Customer requests refund for order #12345 after 14 days."
expected_schema:
type: "object"
required: ["eligible", "refund_amount", "reason_code"]
forbidden_terms: ["I think", "maybe", "as an AI"]
Running Evaluations via CLI
# Run complete evaluation harness against candidate checkpoint moro eval run --checkpoint ./runs/run_01/checkpoints/best/ --suite ./eval/eval_suite.yaml # Compare candidate checkpoint against base foundation model moro eval compare --base Qwen/Qwen2.5-1.5B --candidate ./runs/run_01/checkpoints/best/