CLI Command Group
moro data
Manage dataset ingestion, normalization, information-theoretic deduplication, MI Guard protection, and quality diagnostics.
Subcommands
1. `moro data build`
Executes the complete epistemic curation pipeline on raw files. Automatically redacts PII, scores samples with PPMI and Resnik IC, applies the MI Guard, and produces clean train/val/eval splits.
moro data build [OPTIONS]
| Flag | Type | Default | Description |
|---|---|---|---|
| --source, -s | Path | data/raw/ | Raw input file (.jsonl, .csv, .parquet) or directory. |
| --mi-guard-threshold | Float | 0.85 | PPMI information density threshold to preserve outliers. |
| --dedup-threshold | Float | 0.88 | Cosine similarity threshold for near-duplicate pruning. |
| --pii-scan / --no-pii | Boolean | true | Toggle automatic redaction of emails, phones, and API keys. |
| --val-ratio | Float | 0.10 | Fraction of dataset allocated to validation set. |
2. `moro data report`
Generates an in-depth epistemic report of token entropy, PPMI distribution, and MI Guard statistics:
moro data report --compiled-dir data/compiled/