Epistemic Data Compiler & MI Guard
Information-theoretic dataset curation engineered to mathematically protect critical domain edge cases while eliminating semantic redundancy.
The Fundamental Dilemma in Domain Fine-Tuning
When adapting a large language model to an enterprise domain (such as customer support, legal contracts, or medical documentation), raw input data exhibits extreme semantic skew:
- Boilerplate & Generic Patterns: Greetings, standard policy disclaimers, conversational pleasantries, and common phrases occur thousands of times across the dataset.
- Critical Domain Outliers: Rare legal exceptions, edge-case troubleshooting fixes, specialized product SKU quirks, and safety caveats occur only once or twice.
Traditional deduplication systems (such as MinHash or embedding cosine distance thresholds) rely on cluster density. When they prune datasets to eliminate repetition, high-density generic clusters survive, whereas low-density domain landmarks are discarded as "unrepresentative outliers." The resulting model learns to repeat pleasantries while failing at specialized domain tasks.
The Mutual Information (MI) Guard
MoroAI solves this structural dilemma by introducing the Epistemic Data Compiler with an integrated Mutual Information (MI) Guard. Instead of judging samples solely by semantic similarity, the compiler evaluates their information-theoretic contribution to the domain corpus.
# 1. Calculate Pointwise Mutual Information for token pairs (w1, w2)
PPMI(w1, w2) = max( 0, log2( P(w1, w2) / (P(w1) * P(w2)) ) )
# 2. Compute Resnik Information Content against domain ontology
IC(concept) = -log( P(concept) )
# 3. MI Guard Decision Policy:
if sample.is_cluster_outlier() and sample.mean_ppmi >= mi_guard_threshold:
preserve_sample(sample, category="MI_GUARD_LANDMARK")
elif sample.is_cluster_outlier() and sample.mean_ppmi < mi_guard_threshold:
prune_sample(sample, category="IRRELEVANT_NOISE")
else:
dedup_sample(sample, threshold=dedup_threshold)
Three-Stage Compilation Pipeline
- Format Ingestion & Normalization: Ingests JSONL, CSV, Parquet, or raw text directories. Messages are mapped into a unified schema
{role: "system"|"user"|"assistant", content: string}with automatic schema validation and PII redaction. - Information Density Auditing: Calculates token entropy $H(X) = -\sum P(x) \log_2 P(x)$. Samples with repetitive token spam (e.g. garbled text or low-entropy loops) are flagged for removal.
- Boundary Clustering & Partitioning: Samples are embedded using a local MiniLM/Qwen embedding model. Cosine similarity determines deduplication clusters, while the MI Guard tags protected boundary samples before performing deterministic 80/10/10 train/validation/test splits.
CLI Usage & Pipeline Flags
Compile raw conversational logs into a hardened, balanced dataset ready for training:
# Ingest and compile dataset with MI Guard enabled
moro data build \
--source ./data/customer_support_raw.jsonl \
--mi-guard-threshold 0.85 \
--dedup-threshold 0.88 \
--output ./data/compiled/
# Inspect information-theoretic balance and token distribution
moro data report --compiled-dir ./data/compiled/
Comparison: Naive Deduplication vs MoroAI MI Guard
| Metric / Capability | Naive MinHash / Cosine | MoroAI Epistemic Compiler |
|---|---|---|
| Rare Domain Fact Retention | ~18% (Discarded as outliers) | >99.2% (MI Guard Protected) |
| Boilerplate Pruning | Partial (Threshold dependent) | Comprehensive Entropy Gated |
| Cryptographic Lineage | None | SHA-256 Digest in SBOM |
| PII Pre-Scanning | Manual / External | Built-In Automatic Regex + NLP |