Notation Sets, Living Lexicon, and Cross-Domain Semantic Index
Thenotation-sets proposal establishes a canonical, continuously evolving notation vocabulary supporting Grimoire compression (#309/#182) and cross-domain semantic indexing. It separates canonical notation from domain aliases and domain-specific syntax, connects the vocabulary to the repository’s existing pointer/concept/tool/context indexes, and defines a research-driven evolution loop so the glossary stays synchronized with implementation and operator governance.
Status
- Proposal status:
posted(not accepted) — execution remains gated byrepo-gate+termux-smoke. - Malleable by design: revisions are additive and attributed via the evolution ledger (NSE-004 loop, NSE-008 ledger). This is a revision pass, not a freeze.
- Latest revision: PR #324 — integrated language-density + ADLM + grimoire research (NSE-009 → NSE-018).
Research basis
This revision integrates two prior research sessions:- Language information density & ADLM — Coupé et al. 2019 (cross-language convergence near ~39 bits/s); Petrov et al. EMNLP 2023 (tokenizer inequity across scripts); FLORES-200; tiktoken; Adaptive Dynamic Language Mixing as a tokenizer- and model-specific semantic codec with a canonical-IR truth layer.
- Grimoire compression & category-theoretic IR — layered compression stack (zstd, LZ4, Brotli, FastCDC, xdelta3, CBOR, Parquet); canonical IR for category theory; per-corpus dictionary training.
Item index
Related issues
#320, #309, #182, #175, #126, #304, #196, #177, #208, #274.
Source
Full proposal:docs/proposals/active/notation-sets/ (ITEMS.md, MANIFEST.md) and registry entry in docs/proposals/registry.yaml.