NADA at a glance
How much code goes in, how little essential vocabulary comes out, and how far it reaches.
543
programming languages scanned
of 689 in scope
22,336
library packages mined
across 59 ecosystems
98,522,628
identifiers harvested
every name in every scanned surface
2,552,494
morphemes extracted
379,456 shared across ≥2 ecosystems
15,944
unique terms in the sense graph
the reviewable core vocabulary
190
human locales targeted
every term localizes into each
From raw code to a shared vocabulary
98,522,628
identifiers
→
2,552,494
morphemes (38× smaller)
→
379,456
shared word-parts
What is "the brain"?
NADA distills the world's code down to its morphemes — the reusable word-parts that keywords and library names are built from (head, get, async). These are grouped into collision-firewall clusters of kin terms, then localized together into every human locale so look-alikes stay distinct and synonyms stay consistent. The compiled result — one deterministic, reversible, human-reviewed dictionary that maps every code token to its form in each language — is the brain. It's the supply that lets anyone code in their own language.
Every number here is computed from the live NADA pipeline — nothing is hand-typed.