Full Architectural Specifications

Under the Hood of MoroAI

A deep, uncompromising engineering breakdown of how MoroAI transforms uncurated enterprise data into production-ready, private local language models without cloud dependencies or out-of-memory crashes.

Subsystem 01

Data Compiler & MI Guard

Standard LLM fine-tuning datasets are plagued by two structural flaws: semantic redundancy that dilutes fine-tuning gradients, and the unintended erasure of critical domain facts. When traditional deduplication algorithms (e.g., naive MinHash or embedding cosine cutoff) run, high-frequency boilerplate survives while rare domain edge cases are pruned as unrepresentative outliers.

MoroAI introduces the Mutual Information (MI) Guard. It calculates Shannon entropy $H(X)$ and Pointwise Mutual Information (PPMI) across token n-grams and dense embeddings. Samples possessing high information gain relative to the background corpus are mathematically protected from deduplication pruning:

MI Guard Decision Rule
PPMI(w1, w2) = max(0, log2( P(w1, w2) / (P(w1) * P(w2)) ))
if sample.is_outlier() and sample.ppmi > mi_guard_threshold:
preserve_sample(sample, tag="MI_GUARD_LANDMARK")

Key Capabilities

Positive Pointwise Mutual Information
Cosine + Jaccard boundary clustering
Deterministic train/val/eval splits
Immutable SHA-256 dataset fingerprint
CLI Invocation moro data build
$ moro data build \
    --source ./customer_ops.jsonl \
    --mi-guard-threshold 0.85 \
    --dedup-threshold 0.88 \
    --output ./data/compiled/

[INFO] Total raw inputs: 10,480
[MI-GUARD] Protected 312 critical domain landmarks
[DEDUP] Pruned 2,890 redundant instances
[DONE] Final curated: 7,590 samples
[ENTROPY] +28.4% mean token information gain
[HASH] sha256: 8f3b20e5c941a877028b
              
Subsystem 02

VRAM Recipe Predictor & Hardware Engine

In standard ML workflows, selecting training hyperparameters (micro-batch size, sequence context length, LoRA rank $r$, and gradient accumulation steps) is guesswork. An engineer waits 45 minutes only for the process to terminate with a fatal CUDA out of memory error.

MoroAI introduces the Pre-Flight VRAM Auditing Engine. Prior to launching backpropagation, it profiles the exact GPU architecture (VRAM capacity, memory bandwidth, tensor core generation) and calculates exact peak memory footprint:

Analytical VRAM Formula
M_peak = M_base(NF4) + M_lora(rank, α) + M_opt(AdamW8b) + M_act(seq_len, batch) + M_cuda_ctx

Key Capabilities

NVIDIA CUDA & Apple Silicon Metal detection
Exact activation memory tensor projection
Auto-tuned micro-batch & accumulation
Reproducible config generation (moro.yaml)
VRAM Simulator moro recipe generate
$ moro recipe generate --model Qwen2.5-1.5B

Device: NVIDIA GeForce RTX 4070 (12,288 MB)
Model Weights (4-bit NF4):    ~1,240 MB
Optimizer States (AdamW 8b):   ~450 MB
Activation Peak (Seq 2048):    ~5,120 MB
CUDA Context & Buffers:        ~1,040 MB
-----------------------------------------
Predicted Peak VRAM:           ~7,850 MB
Safe Available Headroom:       ~4,438 MB (36%)

Recipe generated: ./moro.yaml (Verified Zero OOM)
              
Subsystem 03

OOM Auto-Recovery & Self-Healing Training

Unattended training runs fail for two reasons: unexpected CUDA Out-Of-Memory spikes caused by burst sequence outlier batches, or numerical instability causing gradient exploding (NaN/Inf).

MoroAI wraps the backpropagation execution loop in an Autonomous 5-Level Self-Healing Escalation Ladder that intercepts exceptions and resolves them in-flight without aborting the run:

Level 1: Intercept CUDA OOM & flush driver cache + garbage collection
~0.5s
Level 2: Halve micro-batch size and double gradient accumulation steps
Gradient Equiv.
Level 3: Enforce full gradient checkpointing on remaining attention layers
-60% Act. Mem
Level 4: Dynamic sequence length clamping on outlier spike batch
Batch-specific
Level 5: Auto-revert to last stable checkpoint if loss diverges >3x
Rollback
Autonomous Recovery moro train --auto-heal
Step 420/1200 | Loss: 1.12 | VRAM: 7.9GB
[WARN] CUDA OOM at step 421 (token length 3,840)
[HEAL] Level 1: Flushed CUDA memory pool
[HEAL] Level 2: Micro-batch 2 -> 1, Accum 8 -> 16
[RESUME] Checkpoint 400 reloaded
Step 421/1200 | Loss: 1.11 | VRAM: 6.1GB (STABLE)
Step 422/1200 | Loss: 1.09 | VRAM: 6.2GB
[SUCCESS] Full training finished with 0 data loss
              
Subsystem 04

Multi-Layered Evaluation Harness

A low validation loss does not guarantee that a fine-tuned model won't hallucinate or output syntactically broken responses in production.

MoroAI runs four rigorous evaluation barriers before issuing a cryptographic release certification:

  • • Deterministic Regex & Schema Gates: Validates JSON schema integrity, required keywords, and strict negative constraints.
  • • Semantic LLM Judge: Performs reference-based chain-of-thought accuracy evaluation against ground truth domain benchmarks.
  • • Wasserstein Drift Detection: Quantifies output embedding distribution drift against the base model to prevent catastrophic forgetting.
  • • Adversarial Typo Stress Testing: Injects synthetic keyboard typos, homoglyphs, and syntax swaps to test invariance under noise.
Evaluation Gate moro eval compare
Eval Suite: Enterprise Support v1 (150 tests)
-----------------------------------------------
1. Deterministic Format Rules:  100.0%  (PASS)
2. Semantic Domain Accuracy:     94.2%  (PASS)
3. Wasserstein Drift Score:      0.024  (PASS <0.05)
4. Adversarial Typo Stress:      91.8%  (PASS >90%)
-----------------------------------------------
OVERALL VERDICT: QUALIFIED FOR PRODUCTION
Release Gate ID: gate_78f1a9e3
              
Subsystem 05

Closed-Loop DPO Feedback Flywheel

A deployed private model should not remain static. As operators and employees use the model locally, human operators correct errors, provide thumbs-up feedback, and supply manual edits.

The MoroAI Flywheel automatically parses interaction telemetry and extracts Direct Preference Optimization (DPO) preference pairs (prompt, chosen, rejected). When sufficient preference data accumulates, MoroAI triggers a local DPO alignment round:

DPO Objective Function
L_DPO = -E [ log σ( β * log( π_θ(y_w|x) / π_ref(y_w|x) ) - β * log( π_θ(y_l|x) / π_ref(y_l|x) ) ) ]
DPO Miner moro flywheel process
$ moro flywheel process \
    --interactions ./logs/chat_ops.jsonl \
    --min-margin 0.3

[SCAN] Ingested 1,200 production sessions
[PAIR] Extracted 342 chosen/rejected pairs
[FILTER] Discarded 48 ambiguous feedback events
[READY] DPO dataset created: ./dpo_v2.jsonl
[TRIGGER] Scheduled offline alignment run
              
Subsystem 06

Mission Control UI & Real-Time Orchestration

Not every team wants to live solely in the command line. MoroAI includes an interactive, browser-based Mission Control UI that launches with a single command (moro dashboard) and binds to http://localhost:3000.

Built with a zero-cloud architecture using React, Tailwind CSS, and Server-Sent Events (SSE), Mission Control gives operators a single glass pane to orchestrate the entire lifecycle:

Live VRAM & Loss Telemetry: Real-time streaming charts with step-by-step GPU memory gauges.
Visual Recipe Generator: Drag sliders for sequence length and rank with live VRAM safety simulation.
Eval Comparator: Side-by-side completion diffing between base and adapted model checkpoints.
1-Click Ollama Deploy: Export and register models straight into local Ollama runtime.
Start Mission Control moro dashboard
$ moro dashboard --port 3000

╭──────────────────────────────────────────────╮
│  MoroAI Mission Control Dashboard v0.1.0    │
│  Serving on: http://localhost:3000           │
│  Hardware: NVIDIA RTX 4070 (12GB VRAM)       │
│  Active Runs: 1 (run_01_support_qwen)        │
│  Telemetry: SSE connected (120 msg/sec)      │
╰──────────────────────────────────────────────╯
[HTTP] UI bundle loaded in 14ms
[WS] Telemetry pipe active on /api/v1/stream
              
Subsystem 07

Cryptographic Release Governance (SBOM)

In high-compliance environments (healthcare, banking, defense), deploying an AI model without strict provenance is a regulatory non-starter. MoroAI automatically compiles an immutable Software Bill of Materials (SBOM) for every release candidate.

Dataset Hash SHA-256 fingerprint of the compiled training samples
Recipe Snapshot Deterministic seed, learning rate, rank, and target modules
Eval Certificate Signed benchmark scorecard with passing gate thresholds
Weights Signature Cryptographic hash of the output LoRA adapter and GGUF quant
Release Manifest moro release verify
{
  "release_id": "rel_2026_09_28_01",
  "base_model": "Qwen/Qwen2.5-1.5B",
  "dataset_sha256": "8f3b20...a19c",
  "recipe_sha256": "44ea10...e7b2",
  "gate_passed": true,
  "eval_score": 0.942,
  "quantization": "Q4_K_M",
  "provenance_sig": "sig_rsa4096_d892a01",
  "status": "APPROVED_FOR_DEPLOY"
}
              
Subsystem 08

Local Serving & One-Command Deployment

Adapting a model is only half the battle—serving it with ultra-low latency without complex cloud orchestration is where teams get bogged down.

MoroAI automatically fuses trained adapter weights, quantizes to GGUF format (Q4_K_M, Q8_0), generates an optimized Modelfile, and registers the model directly into local runtimes:

Ollama Instant local CLI & REST API
vLLM High-throughput PagedAttention
Docker Isolated sovereign microservices
Deploy Command moro deploy
$ moro deploy --target ollama --name support-v1

[1/3] Fusing LoRA adapter weights...
[2/3] Quantizing to GGUF Q4_K_M (1.1GB)...
[3/3] Registering Modelfile into Ollama...

[SUCCESS] Model registered as: support-v1
Serve locally with:
  ollama run support-v1
Or query via OpenAI-compatible endpoint:
  http://localhost:11434/v1/chat/completions
              
Subsystem 09

Sovereign Privacy & Zero-Telemetry Audit

In an era where third-party AI APIs routinely log queries for model retraining, MoroAI operates under a strict air-gapped sovereign guarantee:

  • Zero Cloud Outbound: No telemetry, error reporting, training tokens, or model weights ever leave your workstation or on-premise VPC.
  • Embedded SQLite: All state, metadata, and run lineage is persisted exclusively in local SQLite databases (.moro/moro.db).
  • Automated PII Sanitization: Built-in pre-training privacy scanner automatically redacts emails, phone numbers, and API tokens before tokenization.
Privacy Audit moro doctor --privacy
$ moro doctor --privacy

[AUDIT] Checking outbound network egress...
  -> Egress sockets: 0
  -> Cloud telemetry: DISABLED
  -> Model cache: /Users/mac/.cache/moro
  -> PII Scanner: ENABLED (Regex + Entity)
  -> SQLite Database: .moro/moro.db (Encrypted)

[VERDICT] 100% AIR-GAPPED SOVEREIGN ENVIRONMENT