Semantic Systems / Language / Glyphs
Strategic Report: Embedded Semantics as a Governed Semantic Identity Layer
Report summary
Research scope and evidence posture. This assessment treats Embedded Semantics as an early technical hypothesis, not an established standard or proven product. The analysis is based on the public surfaces of EmbeddedSemantics.com, Protocol5.com, JustAnIota.com, and MikeKappel.com, together with curr
Key topics
- Semantic Systems / Language / Glyphs
- Semantic Systems
- Language
- Glyphs
- AI
- UAI
- AI Memory
- Agentic Web
- .NET
Research provenance
For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.
Source availability: 116 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.
This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.
Full report
On this page
Executive assessment
Research scope and evidence posture. This assessment treats Embedded Semantics as an early technical hypothesis, not an established standard or proven product. The analysis is based on the public surfaces of EmbeddedSemantics.com, Protocol5.com, JustAnIota.com, and MikeKappel.com, together with current standards and infrastructure documentation available as of August 23, 2026. No source code was assumed or inspected.
The most important empirical fact is that the public Embedded Semantics implementation is presently pre-adoption and effectively pre-registry. Its status page reports zero published concepts, zero reviewed exact expressions, zero represented languages, zero active embedding profiles, and an inactive arbitrary-query semantic runtime. It says an eleven-concept/fifty-five-expression bootstrap bundle exists in the underlying repository but is not deployed to the public registry. The public Concepts page likewise shows zero published concepts.
That matters strategically. The idea should currently be evaluated on architecture, prior art, plausible utility, and testable hypotheses—not on the assumption that market validation, interoperability, or ecosystem effects already exist.
Bottom-line judgment: conditional go, but narrow the thesis substantially.
There is a real and recurring systems problem behind Embedded Semantics, but “stable identifiers for meaning” is not itself a novel category. W3C SKOS already identifies concepts with URIs, attaches multilingual lexical labels, and supports mappings between schemes. UMLS assigns Concept Unique Identifiers to meanings while distinguishing them from strings. HL7 FHIR code systems combine canonical identifiers, versions, concepts, definitions, designations, and governance. ISO 20022 maintains uniquely identified business concepts in a centrally governed dictionary.
That prior art simultaneously validates the architectural pattern and weakens the novelty claim.
The opportunity therefore is not:
“Invent a universal identifier for meaning.”
It is much narrower and potentially more useful:
Create a lightweight, open, federated semantic-identity and binding layer that lets APIs, events, AI agents, AI memory, analytics, and enterprise schemas refer to governed concepts without forcing every participant into the same schema, ontology, knowledge graph, model, or vocabulary.
A good shorthand for the developer problem is semantic foreign keys. A database foreign key preserves reference to an entity independently of its display label; a semantic foreign key would preserve reference to a governed concept independently of the local field name, enum spelling, natural-language label, schema version, or embedding model.
That framing produces a much sharper strategic position:
local representation → semantic binding → stable concept identity → governed record → mappings to other authorities
The stable identifier is not the semantics by itself. The governed record, authority, lifecycle policy, provenance, and mappings make the identifier meaningful.
What the four investigated sites actually establish
| Public surface | What it contributes today | Strategic interpretation |
|---|---|---|
| Embedded Semantics | Separates stable ConceptCodes from language-specific expressions and from embedding/model evidence; exact reviewed lookup is the stated production path, while unseen-query semantic resolution remains experimental. | The cleanest expression of the core architectural hypothesis. It is not yet evidence of scale or adoption. |
| JustAnIota | Publishes draft registry, validation, canonicalization and compact-message concepts; explicitly says Unicode/PUA does not create semantics and that registry records do. Its registry design calls for stable IDs, definitions, examples, policy notes, signing, versioning and immutability after assignment. | Useful adjacent design work around registry discipline and inspectability. |
| Protocol5 | Exposes an evidence workbench and expression/concept experiments and distributes implementation material associated with IOTA/UAI work; its own pages distinguish approximate evidence from exact semantic authority and point to other sites as specification authorities. | Better viewed as an experimental/implementation surface than as independent market validation. |
| MikeKappel.com | Publicly describes related work in persistent AI memory, multi-agent identity, scoped coordination, reviewed memory and handoff architecture. | Provides useful design provenance and shows the broader problem space motivating the concept, but is not independent adoption evidence. |
The governance evidence is especially important. JustAnIota currently describes itself as a single-publisher public record and explicitly says plural governance, broader maintainer rosters, public issue infrastructure and certification are future work. That is an appropriate honest boundary, but it means the surrounding ecosystem should not yet be presented as an external standard-setting community.
Executive answers to the key strategic questions
| Question | Assessment |
|---|---|
| Sharpest problem statement | Systems preserve syntax better than intent. When schemas, names, languages, models or vendors change, the same business concept is repeatedly rediscovered and remapped because there is no lightweight cross-representation identity layer. |
| Who feels it? | Enterprise integration teams, API/platform teams, event-stream owners, agent-platform developers, data-governance teams, analytics/observability teams, and developers maintaining long-lived AI memory. |
| Who adopts first? | AI-agent infrastructure and enterprise platform/integration teams with many independently evolving APIs/events and no dominant domain terminology standard. |
| Most compelling initial uses | Agent capability semantics, API/event semantic bindings, enterprise schema mapping, LLM output normalization, and durable AI memory. |
| Where existing standards are superior | Healthcare primary terminology, financial messaging semantics, industrial information models, government exchange models, observability conventions, localization interchange, and major e-commerce identifiers. Embedded Semantics should map to these rather than replace them. |
| Public registry? | Yes—but as a federated directory/commons and mapping network, not a single universal ontology. |
| Private registries? | Absolutely. Private registries are probably necessary for enterprise adoption. |
| Primary network effect | Reusable mappings and semantic bindings, not merely the number of IDs. |
| What should be open? | Identifier model, registry record model, lifecycle semantics, federation protocol, mapping relations, resolver interface, conformance tests, reference implementation and public registry data that can legally be open. |
| What can be commercial? | Managed registries, governance workflow, change-impact tooling, mapping automation, connectors, enterprise federation, audit/evidence tooling, SLAs and support. |
| What must not become a tollbooth? | Public identifier resolution, specification access, basic compatibility tests, export/self-hosting, public namespace portability, or the ability to use public IDs. |
| Standardize now? | No. The W3C process itself emphasizes wide review and adequate implementation experience before mature standardization. Embedded Semantics currently has neither public registry depth nor independent implementations. |
The recommended strategic posture is therefore:
Build an open semantic identity profile and useful tooling first. Prove that attaching semantic identities measurably reduces integration and migration cost. Let standardization be an outcome of adoption rather than a substitute for it.
Category, problem and alternatives
Recommended category definition: governed semantic identifier layer.
A governed semantic identifier layer is a thin interoperability layer in which a stable identifier refers to a versioned, governed concept record, while local strings, enum members, schema fields, natural-language expressions, model outputs and external terminology codes can be bound or mapped to that concept.
Four separations are essential:
- Identity is not representation.
customer,cliente,顧客,cust_id, or an embedding vector can be evidence about a concept; none should automatically become the concept's durable identity. - Identity is not resolution. Deterministic lookup or probabilistic retrieval finds an identifier. It does not create the identifier's authority.
- Identity is not ontology. An atomic identifier can state “this field refers to concept X” without requiring a complete logical model of X's relationships.
- Identity is not entity identity. A Concept ID for “customer account” is different from the identifier of customer account instance
#875421.
Embedded Semantics already makes the particularly sound architectural choice of separating model vectors from registry authority and of abstaining when its exact registry lacks sufficient evidence. Its public methodology states that unseen-expression retrieval, reranking and confidence calibration remain experimental and that the target is concept resolution rather than generic topic similarity.
The sharp problem
Modern software has become good at exchanging structure while remaining surprisingly fragile about intent.
A schema can say:
field: customer_status
type: string
allowed: ["A", "I", "S"]
but another system may have:
accountState: ["active", "disabled", "suspended"]
and an agent may expose:
tool: retrieve_active_relationships
A human can often infer overlaps. A schema validator cannot determine whether A and active express exactly the same governed business state, whether I means “inactive” or “invalid,” or whether “customer” and “account holder” are equivalent in the relevant business process.
Companies therefore accumulate mapping tables, ETL transforms, prompt instructions, glossary documentation, one-off agent descriptions and tribal knowledge.
The potentially valuable primitive is:
representation A ─┐
representation B ─┼──> stable governed concept ID
language C ─┤
agent description ┘
Then the local representation can change without automatically invalidating the semantic reference.
The strongest problem statement is consequently:
Distributed systems need a lightweight way to preserve references to intended concepts across representation changes—schema versions, renamings, languages, vendors and model generations—without requiring every participant to share a complete ontology or a single physical schema.
The phrase “lightweight” is decisive. Remove it and the problem has already been attacked for decades by RDF, ontologies, master-data systems, terminology servers, knowledge graphs and enterprise semantic layers.
Why the idea is plausible but not novel
SKOS is the closest open conceptual prior art. It already supports URI-identified concepts, Unicode lexical labels in multiple languages, preferred and alternative labels, definitions, hierarchical/associative relationships and exact/close mappings to concepts in other schemes.
UMLS is even closer to the “strings are not meaning” thesis. The National Library of Medicine explains that a concept is a meaning, that one meaning can have multiple names, and that its Metathesaurus separates Concept Unique Identifiers from lexical, string and atom identifiers.
FHIR similarly treats a code system as a governed definition of codes and meanings, with a globally stable canonical URL and independent version information.
ISO 20022's Data Dictionary identifies business concepts independently of the messages in which they occur, and its Business Model changes are handled through explicit change requests to a Registration Authority.
The conclusion is not that Embedded Semantics is unnecessary. It is that its potential innovation is operational packaging and cross-domain interoperability, not the philosophical discovery of stable concepts.
Alternative-technology comparison
| Alternative | What it does very well | Where it fails or becomes awkward for this problem | Does a governed semantic-ID layer add something? |
|---|---|---|---|
| Strings | Human readability; nearly zero implementation cost. | Renames, synonyms, localization, casing and divergent terminology make equality unreliable. | Yes, whenever representation changes but concept identity should not. |
| Enums | Strong local compile-time/runtime constraints. | Identity usually belongs to one codebase/schema; equivalent enums in other systems still require mappings. | Yes, as a cross-system identity above local enums. |
| JSON Schema / Protobuf / Avro / API schemas | Structural contracts, typing and compatibility. Confluent Schema Registry, for example, explicitly manages schema evolution and producer/consumer compatibility. | Structural compatibility does not by itself prove two differently named fields or values have the same business intent. | Frequently, provided the semantic layer remains supplemental rather than replacing schema validation. |
| Traditional taxonomies / SKOS | Stable concepts, multilingual labels, hierarchy and mappings already exist. | Developer ergonomics and runtime binding to APIs/events/agents are not inherent deployment outcomes of SKOS itself. | Only if the new system is substantially easier to bind into software and interoperates with SKOS rather than reinventing it. |
| Formal ontologies / RDF | Rich relationships, explicit graph semantics, reasoning and globally scoped IRIs. RDF 1.2 uses absolute IRIs as identifiers in its abstract model. | Often more modeling and infrastructure than a developer needs merely to say “these two fields refer to the same concept.” | Yes as an ontology-light profile; no if it starts rebuilding OWL/RDF badly. |
| Domain code systems | Excellent when a legitimate authority already governs the domain. FHIR and UMLS show mature implementations of this pattern. | Different authorities and industries still need mappings; many internal enterprise concepts have no public code system. | Yes as federation/crosswalk infrastructure. It should not mint competing IDs for already-governed concepts without cause. |
| Embeddings | Flexible approximate semantic retrieval, paraphrase matching and discovery. | Their coordinates and neighborhoods are model-dependent; proximity does not itself constitute a governance decision that two concepts are equivalent. Embedded Semantics correctly treats vectors as evidence rather than authority. | Strongly complementary: use embeddings to find candidate IDs, not define them. |
| Vector databases | Efficient retrieval over embeddings and unstructured material. | They optimize finding neighbors, not governing durable public concept identity. | Yes, if the ID layer becomes the durable output of retrieval. |
| Knowledge graphs | Rich entity/concept relationships, graph traversal and integration. | A graph is a much larger architectural commitment than attaching a concept identity to an API field or event attribute. | Sometimes. Semantic IDs can be inputs to or references from a KG. |
| LLM prompting | Extremely flexible interpretation and transformation without pre-modeling every mapping. | A prompt is not a durable external authority, and repeated interpretation can vary with context/model/runtime changes. | Yes where deterministic, auditable identity matters after the LLM has interpreted an input. |
| Enterprise data catalogs / glossaries | Already attach governed terms to enterprise assets. DataHub explicitly separates stable glossary identifiers from human-friendly display names; OpenMetadata supports glossary concepts, synonyms and relationships. | Usually optimized around one organization's catalog rather than neutral cross-organization protocol-level interchange. | Potentially—but integration with these products should be a first-class feature, not competition by default. |
| Enterprise ontologies | Can model an organization's full operational world. Palantir's Foundry Ontology, for example, maps data into object, property, link and action types. | Much broader than an open identifier primitive and tied to an operational platform. | Yes if the new layer is portable, small and vendor-neutral; otherwise it becomes another ontology platform. |
Where existing technologies should win
A skeptical strategy must explicitly concede domains where a new semantic authority would be counterproductive.
Healthcare: do not mint a parallel concept for “headache” merely because the system can. UMLS and clinical code systems already exist precisely to normalize meanings across terminology sources, while FHIR provides mature code-system/value-set machinery. The useful role is mapping, provenance, federation and developer binding around those authorities.
Finance: ISO 20022 has a governed, syntax-independent Data Dictionary containing uniquely identified business concepts and formal maintenance procedures. A new generic registry should consume or reference that authority rather than compete with it.
Industrial systems: OPC UA Companion Specifications explicitly exist to provide domain information models and semantic interoperability, including models for robotics. The OPC Foundation recommends reusing existing Companion Specifications to increase interoperability.
Observability: OpenTelemetry Semantic Conventions already define common attribute names, types, meanings and valid values, with formal stability levels and code ownership requirements.
Government exchange: NIEM already targets syntactic and semantic interoperability across public-sector domains and supports representations such as XML, JSON and OWL.
In all these cases, the potential Embedded Semantics product is a bridge, adapter and durable crosswalk layer, not a replacement vocabulary.
That distinction is strategically critical.
Competitive landscape and use-case portfolio
The relevant competitive landscape is much broader than “semantic databases.” Embedded Semantics would sit in the overlap between several established categories.
Concept/vocabulary standards: SKOS/RDF provide open Web primitives. UMLS/FHIR demonstrate mature clinical terminology identity. ISO 20022 demonstrates governed business concepts. OPC UA Companion Specifications demonstrate industrial semantic models. NIEM does so for government. OpenTelemetry does so for telemetry. These systems demonstrate that semantic identity is valuable when backed by an authoritative community.
Schema registries and data contracts: Confluent Schema Registry solves producer/consumer schema evolution and compatibility. That is adjacent rather than equivalent: semantic identity should complement the structural contract.
Enterprise catalog and glossary platforms: DataHub and OpenMetadata already treat business terms as governed semantic objects attachable to datasets and fields. DataHub explicitly recommends stable glossary-term identifiers separate from mutable display names.
Enterprise ontology platforms: systems such as Palantir Foundry Ontology provide much richer operational modeling. An Embedded Semantics layer must remain dramatically simpler and more portable to avoid competing head-on with those architectures.
AI interoperability protocols: MCP currently identifies tools by a name and schema; its July 28, 2026 specification defines the open protocol for integrating LLM applications with external tools and data. A2A addresses communication and collaboration between agents and has been moved into open Linux Foundation stewardship. Neither fact establishes a universal business-concept identity system. That leaves a plausible extension point rather than a reason to create a competing agent protocol.
The strongest white space is consequently:
Portable semantic bindings between existing contracts and existing authorities.
That is a much more defensible niche than “the semantic standard.”
Top use cases and three-dimensional ranking
The following ranks are strategic estimates, not measured market-size statistics. “Technical value” means incremental value supplied by stable governed identity relative to the strongest existing alternative. “Adoption probability” considers integration friction, current standards gaps and buyer/developer incentives. “Market value” considers likely economic importance of the underlying integration problem rather than near-term revenue certainty.
| Technical rank | Use case | Adoption rank | Market rank | Strategic verdict |
|---|---|---|---|---|
| 1 | Cross-standard concept federation and crosswalks | 13 | 13 | Architecturally the purest problem: an ISO/industry code, enterprise term, API concept and agent concept can remain under separate authorities while mappings connect them. No universal owner is required. This is likely core infrastructure rather than the first monetizable wedge. SKOS already provides useful mapping relations to learn from. |
| 2 | Enterprise schema/data-contract semantic mapping | 6 | 1 | Extremely strong. Companies already maintain evolving schemas and contracts; semantic IDs can sit above them and preserve conceptual mappings as physical names change. Confluent demonstrates the demand for contract/version governance, while DataHub demonstrates demand for stable business terms. |
| 3 | Standards metadata and semantic-crosswalk registry | 16 | 20 | High public-good value: give concepts in many specifications machine-resolvable identities and mappings. Commercial value is weaker unless attached to tooling. |
| 4 | AI-agent capability/tool intent identity | 1 | 6 | Best early wedge. MCP tools currently have locally defined names and schemas; stable semantic metadata could let registries say that differently named tools expose equivalent or related capabilities without changing MCP transport. |
| 5 | Healthcare terminology bridge/provenance | 19 | 7 | High technical consequence, low probability as an initial market. Use existing UMLS/FHIR/SNOMED-like authorities; add bindings and provenance. Do not become a new clinical terminology authority. |
| 6 | Event-bus semantic identities | 4 | 4 | Strong wedge. Producers and consumers can bind event types/fields to concepts independently of topic naming and schema versioning. Schema registries remain responsible for structural compatibility. |
| 7 | Financial message/data semantic bridge | 18 | 2 | Very valuable economically, but ISO 20022 already owns substantial semantic ground. The opportunity is legacy/proprietary-to-ISO mapping and cross-standard binding, not replacement. |
| 8 | Durable AI-memory concept anchors | 5 | 11 | Strong technical hypothesis. Memory entries could refer to stable business concepts while chunks, summaries, embedding models and retrieval indices are rebuilt. The reviewed-memory work described in the investigated project ecosystem makes this a natural internal proving ground, though not external proof. |
| 9 | Analytics metric and dimension identity | 7 | 5 | Valuable where “revenue,” “active customer,” “conversion,” etc. have multiple implementations. IDs could anchor definition and lineage across warehouses/BI tools without forcing one physical metric layer. |
| 10 | API operation/parameter semantic binding | 3 | 3 | Excellent. A concept identifier attached to an API operation, parameter or enum value could survive renaming and allow automated compatibility/mapping tools to reason above local OpenAPI names. |
| 11 | Government interagency semantic mapping | 17 | 18 | Strong technical value but NIEM already covers substantial government semantic/syntactic interoperability. Best as a bridge to NIEM and other public vocabularies. |
| 12 | Industrial digital-twin vocabulary bridge | 15 | 12 | Useful only as federation among OPC UA Companion Specifications, enterprise models and other vocabularies. OPC UA already offers domain semantic interoperability. |
| 13 | LLM structured-output normalization to governed concepts | 2 | 8 | Highly adoptable. Let the model generate candidates, then require resolution to approved IDs or abstention. This cleanly combines probabilistic interpretation with deterministic downstream business logic and matches Embedded Semantics' own authority/evidence separation. |
| 14 | Robotics capability/action semantics | 12 | 16 | Interesting for agentic robotics and heterogeneous fleets, but existing ROS/OPC-domain models reduce the need for another broad authority. OPC UA Robotics already provides standardized cross-vendor information modeling. |
| 15 | IoT capability/property semantic bridge | 14 | 15 | Useful in heterogeneous environments, especially between existing information models; not an obvious first market because industrial standards already address much of the problem. |
| 16 | Multilingual enterprise glossary/search normalization | 9 | 14 | A reasonable application of stable concept identity. SKOS already supports labels in multiple languages, so value must come from easier runtime binding, governance and resolver tooling. |
| 17 | Legal clause/obligation concept references | 20 | 17 | Potentially important but semantically dangerous: legal meaning depends heavily on jurisdiction, instrument, version and context. Existing legal identifiers/document standards should be preserved; only narrowly scoped concept binding should be attempted. |
| 18 | Observability convention aliasing and migration | 8 | 9 | Valuable as migration tooling, but not as a replacement for OpenTelemetry Semantic Conventions, which already govern names, types and meaning and have stability rules. |
| 19 | E-commerce attribute/category normalization | 11 | 10 | Real mapping pain, but GS1 and Schema.org already occupy substantial shared vocabulary/identifier territory. The better role is cross-marketplace mapping and private extensions. |
| 20 | Localization semantic message/intent keys | 10 | 19 | Stable concept keys can avoid source-string identity, but localization standards and mature application message-ID practices already solve much of this. Use only where the identity needs to cross applications/systems, not as a new translation standard. |
The five use cases worth pursuing first
Agent capability identity is the best ecosystem beachhead. MCP and A2A have generated substantial activity around protocol interoperability, but protocol interoperability does not automatically provide common domain semantics. MCP tools are uniquely named within their exposed metadata and carry schemas; an optional semantic identity field could add meaning without replacing MCP. The Linux Foundation's investment in neutral governance around MCP and related agent infrastructure also shows that developers care about vendor-neutral foundations in this layer.
Enterprise API/schema mapping is probably the best commercial beachhead. It addresses an expensive existing integration workflow rather than asking buyers to believe in a new philosophical category. The pitch is not “adopt semantics”; it is “stop remapping the same business concepts every time a service or schema changes.”
Event-bus semantics is the cleanest demonstration of the “semantic foreign key” idea because producer and consumer evolution is inherently decoupled. Existing schema registries solve shape compatibility; concept identity can prove complementary value if it reduces human mapping effort.
LLM normalization exploits a contemporary architectural weakness. An LLM can interpret loose input; the registry can turn accepted interpretation into a deterministic business concept. Embedded Semantics' insistence that uncertainty can result in abstention rather than forced assignment is the right safety principle.
Durable AI memory is a compelling proving ground because embeddings, chunk boundaries, summaries and even model families are replaceable implementation artifacts. A memory architecture that stores stable concept references alongside those artifacts can be tested through forced migrations. This is a hypothesis that should be experimentally demonstrated rather than asserted.
The anti-use cases
The idea should explicitly say no when:
- one schema is controlled by one application and no cross-boundary reuse exists;
- an enum already has a stable domain code and every participant accepts that authority;
- the use case needs rich logical reasoning rather than identity;
- the semantics are so contextual that an atomic concept ID gives false confidence;
- a mature industry standard already defines the exact relevant concepts;
- fuzzy retrieval is sufficient and no durable downstream contract depends on the result;
- the cost of governing concepts exceeds the cost of maintaining the mapping.
This willingness to reject use cases would improve credibility substantially.
Adoption and ecosystem architecture
Ideal early adopters
The best early adopter is not “a company that wants semantics.” It is an organization with visible representational churn around relatively stable business concepts.
| Early-adopter profile | Why it is attractive | What to prove |
|---|---|---|
| AI-agent platform/framework developers | Many tool names, descriptions and schemas are locally defined while agent interoperability is rapidly standardizing around MCP/A2A. | Different implementations can recognize equivalent capabilities without shared tool names. |
| Internal developer-platform teams | They own API catalogs, event catalogs, service metadata and developer tooling and can introduce optional metadata without changing business applications immediately. | Semantic bindings survive API/event renames and reduce integration work. |
| Large integration-heavy enterprises | They already carry mapping debt among ERP, CRM, data warehouse, APIs, events and acquired systems. | Reusable concept bindings materially reduce repeated mappings. |
| Data catalog/governance vendors or sophisticated users | They already manage glossary IDs and relationships. DataHub/OpenMetadata show the category exists. | Semantic IDs become portable outside the catalog into runtime contracts. |
| Analytics/platform teams | Metric and dimension semantics often recur across physical systems. | Stable identity improves lineage/change-impact analysis. |
| Observability vendors/platform teams | Existing semantic conventions provide an authority to map to. | Automatic mapping/migration creates value without competing with OpenTelemetry. |
The wrong first customers are standards committees, hospitals, banks, national governments or industrial automation consortia. Their problems are important, but adoption requires institutional authority, certification, regulatory confidence and compatibility with mature standards. They are later proving grounds, not a startup wedge.
Developer adoption path
The system needs to be adoptable without asking a developer to operate an ontology stack.
The ideal sequence is:
Start as metadata. A developer attaches an optional semantic ID to an existing tool, event field, API element or memory record. Nothing about the transport changes.
For example, conceptually:
{
"name": "customer_status",
"type": "string",
"semantic_id": "https://example.org/concepts/customer-status"
}
The key is not that this JSON syntax becomes normative. The key is that adding semantics must feel like adding an annotation rather than adopting a new application framework.
Resolve in development tools before runtime. A CLI/IDE extension should search approved registries, display definition/authority/mappings and let the developer bind a field deliberately. Probabilistic search can suggest; governance confirms.
Validate in CI. CI should answer: Does the concept exist? Is it deprecated? Has its authority changed? Is this mapping allowed? Does the dependency pin an acceptable registry version? Has a concept split?
Generate language constants. SDK generation should turn IDs into typed constants so application code does not repeatedly manipulate raw URLs/strings.
Make runtime lookup optional. Applications must be able to operate from signed/cached registry snapshots. A public registry outage must never halt a production payment, robot, clinical workflow or agent.
Add probabilistic resolution last. The resolver is a convenience layer, not the category. This ordering matches Embedded Semantics' existing architectural separation between exact registry authority and experimental model-based retrieval.
For MCP specifically, the current specification already provides names, schemas and extensible metadata surfaces, making a semantic-identity extension/profile more plausible than inventing a rival agent protocol.
Enterprise adoption path
An enterprise rollout should be overlay-first rather than migration-first.
First import existing authorities: internal business glossaries, data-catalog terms, ISO/industry identifiers and OpenTelemetry/domain vocabularies. The system should not begin by minting thousands of new concepts.
Then bind concepts to a small set of high-friction integrations. Keep original schemas unchanged. Run the semantic layer in “shadow metadata” mode and measure whether it correctly predicts/reuses mappings during schema changes.
Only after value is demonstrated should governance become enforceable in CI or deployment.
This is where integration with DataHub, OpenMetadata and schema registries matters strategically. DataHub already attaches governed terms to datasets and fields, while Confluent already manages schema compatibility. A successful semantic-ID layer should enrich that ecosystem rather than ask an enterprise to discard it.
Public registry versus private registries
There should be a public registry, but it should not be “the registry of all meaning.”
A monolithic universal registry would immediately create governance disputes: Who gets to define “customer,” “risk,” “resident,” “employee,” “material,” or “AI agent”? Meanings can vary legitimately by jurisdiction, industry and organization.
The better architecture is federated authority:
Public directory / mapping commons
/ | \
/ | \
Standards Vendor Community
namespaces namespaces namespaces
| |
enterprise mappings |
| |
Private registries ---+
The public layer should primarily provide:
- authority/namespace discovery;
- immutable public concept records where an appropriate public maintainer exists;
- public crosswalks;
- signatures and release metadata;
- mirrors/snapshots;
- governance records;
- conformance artifacts.
Enterprises should be able to run private registries using exactly the same protocol and record format.
A company might own:
https://semantics.example.com/finance/net-revenue
while mapping it to public analytics, accounting or industry concepts where appropriate.
Private concepts need not become public merely to participate in federation. This matters for confidential business models and proprietary processes.
Do not invent a new global identifier syntax
One of the highest-leverage strategic choices would be to avoid creating a proprietary ConceptCode naming universe as the only global identity mechanism.
W3C technologies already use globally scoped IRIs, SKOS identifies concepts using URIs, and FHIR uses stable canonical URLs for code systems.
Therefore:
ConceptCodecan remain a useful compact/local alias.- Every externally interoperable concept should have a canonical URI/IRI.
- The URI should identify the authority as well as the concept.
- Dereferencing it should preferably return human and machine representations.
- No semantic inference should depend on parsing the identifier text.
This avoids asking developers and standards organizations to accept a new identifier infrastructure before they can use the semantic model.
The registry record that would actually matter
The durable asset is not the identifier string. It is the governance contract around it.
A minimal record should contain approximately:
canonical identifier
authority / namespace owner
status
definition
scope note
preferred and alternate labels
locale / language metadata
creation metadata
provenance
external mappings
examples
counterexamples / hard negatives where useful
supersession / deprecation history
split / merge relationships
registry snapshot/version
signature / integrity data
JustAnIota's current registry work is directionally consistent with this: it calls for stable IDs, definitions, locale notes, examples, policy information and a production registry that is versioned, signed, immutable after assignment and validator-tested.
One important refinement is needed: the ID should be immutable; the entire record should not necessarily be frozen forever.
A typo correction or added French label should not mint a new semantic identity. A material change to the intended concept should.
The lifecycle rules therefore need explicit distinctions among:
- editorial change;
- added evidence/labels;
- non-breaking clarification;
- deprecation;
- exact replacement;
- merge;
- split;
- material semantic redefinition.
A silent semantic redefinition under the same identifier should be forbidden.
Federation mechanics
Federation should be designed around authority, not consensus on one truth.
Registry A should be able to say:
A:customer-account
closeMatch B:client-account
exactMatch C:party-account-holder
without B or C being forced to adopt A's concept.
SKOS's existing distinction among mapping relationships such as exact and close mappings provides a useful starting point instead of inventing relation names from scratch.
Mappings themselves must have provenance. An exactMatch assertion by Company A is not automatically an assertion by Company B.
A robust federation model therefore makes these separately identifiable:
concept
authority
mapping assertion
mapping author
mapping evidence
mapping status
mapping version
This prevents the registry from becoming a political battle over who is “correct.”
Network effects that could realistically emerge
The strongest potential network effect is mapping liquidity.
If ten organizations independently bind APIs to the same public concepts, an eleventh can reuse more mappings. If two registries publish high-quality crosswalks, integrations between their ecosystems become easier. Each additional high-quality binding potentially increases the reuse value of existing concepts.
A second effect is tooling compatibility. As SDKs, schema registries, MCP servers, API gateways, event catalogs, observability platforms and data catalogs understand the same semantic metadata, every registry becomes more useful.
A third is resolver evidence. Reviewed aliases, multilingual expressions and hard negatives can improve candidate retrieval. Embedded Semantics explicitly treats reviewed positives and hard negatives as evaluation inputs for future resolution.
A fourth is governance reputation. A namespace that has stable stewardship, clean lifecycle history and reliable mappings becomes more trusted over time.
But there is an important negative conclusion: the raw number of concepts is not a strong network effect. A registry with ten million poorly governed IDs is worse than a registry with ten thousand widely reused concepts. Optimizing for concept count would reproduce the worst failure modes of taxonomies.
Governance, standards and business model
What should be open
A credible ecosystem requires a hard separation between shared interoperability infrastructure and commercial operational services.
| Must be open and independently implementable | Reason |
|---|---|
| Identifier and namespace model | Otherwise every semantic reference creates vendor dependence. |
| Concept-record schema | Independent registries must exchange records. |
| Lifecycle/deprecation/merge/split rules | Long-lived references depend on deterministic semantics. |
| Mapping relation vocabulary | Federation depends on interoperable crosswalks. |
| Resolver request/response contract | A company must be able to replace one resolver with another. |
| Registry federation/discovery protocol | No central service should own interoperability. |
| Signature/integrity representation | Trust verification cannot depend on a closed server. |
| Conformance test suite | External implementations need objective compatibility tests. |
| At least one reference registry/server | Developers need a self-hostable baseline. |
| SDKs/CLI for major languages | The metadata needs to become ordinary developer infrastructure. |
| Benchmark/evaluation harness | Resolver quality claims must be independently reproducible. |
| Public registry snapshots/dumps | Public IDs must remain usable if a vendor disappears. |
| Governance/RFC history | Semantic authority needs auditable decision making. |
The open-source code license should be permissive enough for enterprise and vendor adoption. The data license should likewise permit broad reuse where the project actually has rights to license the data; imported mappings to third-party standards must preserve the terms of their original authorities.
What can be commercial
A company can build a significant product around an open semantic layer without owning the semantic primitive.
Commercial opportunities include:
Managed private registries. Multi-tenant hosting, access control, backups, high availability, regional deployment and enterprise support.
Governance workflow. Proposals, approvals, reviewer assignment, policy gates, delegation and audit history.
Mapping workbench. Candidate generation, human review, bulk reconciliation, merge/split workflows and conflict detection.
Semantic CI/CD. Breaking-change detection, API/event checks, deprecated-concept detection and dependency-impact analysis.
Enterprise connectors. DataHub, OpenMetadata, Kafka/schema registries, API gateways, data warehouses, BI layers, MCP catalogs, event catalogs and industry vocabulary imports.
Federation gateway. Private/public namespace routing, caching, signing, policy enforcement and cross-enterprise mappings.
Resolver/evaluation infrastructure. Hosted multilingual candidate retrieval, benchmark management and model evaluation—provided customers can swap it out.
Compliance/audit packages. Evidence that a particular mapping or concept version was in force for a transaction or software release.
Premium curation. Expert-assisted mapping and terminology governance for specialized domains.
Support and certification services. Paid certification can be legitimate if the normative compatibility requirements and self-test suite remain public.
The monetization principle should be:
Sell operation, governance, automation and assurance—not permission to participate in the semantic network.
What should never be monetized
The project should never make any of the following proprietary tollbooths:
- the right to use a public concept identifier;
- basic dereferencing of public identifiers;
- the normative specification;
- the ability to implement a compatible registry;
- the basic conformance test suite;
- export of a customer's own private registry;
- offline resolution of public concepts from snapshots;
- interoperability between private and public namespaces;
- the right to create an independent namespace;
- mandatory access to one company's resolver;
- access to historical public concept versions necessary to interpret old data.
Those restrictions would turn every external reference into a future ransom opportunity and would make serious standards adoption irrational.
How to avoid ecosystem capture
Vendor capture can occur even with “open source” if the vendor controls the only registry, trademark, test suite, hosted discovery service, identifier allocator or release process.
Avoid it structurally:
Namespace delegation, not central allocation. Organizations should be able to use URI namespaces they control.
Multiple implementations from the beginning. The specification should be written so a second server can pass conformance without reading the first implementation's internals.
Portable data. Registry exports should include full lifecycle and provenance records in normative formats.
Mirrors. Public namespaces should be independently cacheable and mirrorable.
Open compatibility marks. Trademark rules can protect misuse of the project name without requiring customers to buy a service.
Neutral technical governance before standardization. A commercial company may remain the leading implementer, but it must no longer possess unilateral control over the interoperability contract.
No proprietary extensions in the critical path. Premium features can exist, but a semantic ID issued in the open layer must not become unusable without them.
Current governance credibility
At present, governance credibility is the largest gap between a technical experiment and an ecosystem.
JustAnIota explicitly documents a single-publisher model and says broader maintainer rosters, a public governance body, public issue tracker/forum/RFC portal and plural governance are future work. Embedded Semantics currently has no deployed public concept corpus from which an external community could even evaluate stewardship behavior over time.
That is acceptable for a research project. It is insufficient for a semantic authority.
External developers would need to see, at minimum:
- public source repositories;
- public issues and design discussions;
- contributor and decision policies;
- a published intellectual-property policy;
- at least several maintainers not controlled by one organization;
- reproducible releases;
- conformance tests;
- transparent security reporting;
- independent implementations;
- stable release and deprecation policy;
- governance over mappings as well as IDs;
- a clear appeal/dispute process for contested concepts;
- public records showing how real merge/split/deprecation disputes were resolved.
OASIS Open Projects offer one possible organizational pattern through project governing boards and maintainers. The Linux Foundation's stewardship of agent infrastructure is another relevant ecosystem precedent. MCP was placed in the Agentic AI Foundation alongside other open agent projects, while A2A was donated into Linux Foundation governance with major multi-vendor participation.
The lesson is not “join a foundation tomorrow.” It is that infrastructure gains credibility when governance is visibly separable from the commercial interests of one publisher.
Standards strategy
Do not seek formal standard status yet.
The W3C process is useful here because it explicitly requires wide review, consensus-building and adequate implementation experience; Candidate Recommendation exists in part to gather evidence that a specification actually works in practice. W3C also now has an explicit Registry Track for standards-associated registries. ISO similarly describes its process as expert, multi-stakeholder and consensus based.
Embedded Semantics is currently several stages before that threshold.
A sensible standards ladder is:
| Maturity | Vehicle | Purpose |
|---|---|---|
| Early | Public RFC/specification repository | Let developers attack the model while it is cheap to change. |
| Pilot | Open-source project + independent implementations | Discover which parts must actually interoperate. |
| Community | Neutral multi-vendor working group/open project | Resolve governance and namespace questions. |
| Domain profile | Existing ecosystem groups | Standardize bindings for MCP/A2A, OpenTelemetry, industry systems, etc., rather than force one giant spec. |
| Mature Web/enterprise layer | W3C or OASIS | Candidate homes if the generic registry/federation model has demonstrated cross-vendor value. |
| International maturity | ISO/IEC where warranted | Only after broad international and multi-industry implementation. |
NIEM Open offers a contemporary precedent for this type of maturation: work developed under OASIS was submitted to ISO/IEC JTC 1 in June 2026.
Standards organizations that could eventually matter
W3C is the strongest candidate if the core becomes Web-native concept identification, mappings and registry federation. Alignment with IRIs, SKOS, RDF/JSON-LD and the W3C Registry Track would dramatically reduce reinvention.
OASIS is attractive if the center of gravity becomes enterprise interoperability, signed registry records, profiles and federation. Its Open Project governance model provides a bridge between code and specifications.
Linux Foundation / Agentic AI Foundation and A2A ecosystem matter if the first successful use is agent capabilities. The correct strategy would be an optional semantic profile or metadata convention around existing agent standards, not a competing transport.
HL7 and clinical terminology authorities matter for healthcare mappings, but should remain semantic authorities for their domains.
ISO 20022 governance bodies matter for finance, again as authorities to map to rather than displace.
OPC Foundation matters for manufacturing, IoT and robotics because its Companion Specifications are already designed around semantic interoperability.
OpenTelemetry matters for observability-specific bindings because it already operates an evolving semantic-convention process.
NIEM/OASIS and eventually ISO/IEC matter for government data.
A new identifier URI scheme should be avoided; therefore the project should not create unnecessary IETF standardization work merely to obtain a branded identifier syntax.
Risks of premature standardization
Premature standardization would likely freeze the wrong abstraction.
The project does not yet know whether external users primarily need concept identity, concept-to-schema bindings, mapping provenance, versioned registry snapshots, probabilistic resolution, namespace federation or some combination. A specification designed before pilots would encode assumptions rather than interoperability requirements.
There are six specific dangers:
Ontology creep. Every adopter asks for richer relationships until the “lightweight” layer becomes another ontology language.
Resolver creep. An implementation-specific embedding pipeline becomes normative even though models will change much faster than identifier governance.
Identifier proliferation. A new universal ID is minted where existing SKOS/FHIR/ISO/OPC identifiers were sufficient.
Governance hardening. Founder-era decision processes become accidentally institutionalized.
Premature compatibility promises. Fields that should change become frozen because they appeared in an early public spec.
Standards theater. Formal-looking documents create credibility signals without the expensive part of standardization: independent implementation, stakeholder review and negotiated consensus.
The current JustAnIota roadmap itself acknowledges a similar principle: planned ideas should not become current support until public artifacts, validator behavior, implementation evidence and release records agree. That discipline should be applied even more strongly to Embedded Semantics.
Vendor-lock-in risks
The most dangerous lock-in is not proprietary source code. It is semantic dependency.
If customers put one vendor's identifiers in ten years of events and records, changing the identifier authority becomes much more difficult than replacing an ordinary SaaS product.
Therefore the project needs an unusually high openness bar.
A semantic infrastructure vendor should assume:
Every identifier it persuades a customer to persist is a long-term obligation to that customer and therefore must outlive the vendor.
That is why public snapshots, namespace portability, federation, neutral governance and complete exports are not nice-to-have open-source gestures. They are prerequisites for responsible adoption.
Proof strategy, messaging and partnerships
Demonstrations that would actually prove the idea
A homepage demo where “hello,” “hola” and “こんにちは” resolve to one concept would not prove much. SKOS-style multilingual labels and decades of terminology systems already establish that such mappings are possible.
The demonstrations need to prove incremental operational value.
| Demonstration | Setup | Competing baseline | Evidence of success |
|---|---|---|---|
| Agent capability interoperability | Three independently implemented agents expose differently named tools with structurally different schemas but overlapping capabilities. Attach semantic IDs and mappings. | MCP names/descriptions/schemas plus LLM reasoning alone. | Higher deterministic capability matching, fewer incorrect tool matches, interoperability after tool renames. |
| Schema migration survival | Version an API/event schema repeatedly: rename fields, split enums, move fields, change language labels. | Ordinary schema registry + manual mappings. | Existing semantic bindings survive non-semantic changes and correctly flag genuine semantic changes. |
| Durable-memory migration | Build memory with embedding model A; later rechunk/re-embed with model B and modify summaries. | Vector/chunk IDs only. | References to governed concepts remain stable, and application behavior remains reproducible despite retrieval-layer replacement. |
| Multilingual hard-negative benchmark | Use several languages with close but non-equivalent terms, ambiguity and culture-specific distinctions. | Embedding nearest-neighbor and LLM prompt baselines. | High precision on accepted mappings, explicit abstention on ambiguity and documented false-neighbor performance. Embedded Semantics already proposes Recall/MRR/false-neighbor/abstention metrics; those now need public datasets and results. |
| Public/private federation | Public registry, two independent enterprise registries and one industry vocabulary. Disconnect the public service after synchronization. | Central hosted lookup. | Offline resolution continues; provenance remains verifiable; different authorities can disagree without corrupting IDs. |
| Concept lifecycle | Force a concept through typo correction, label addition, deprecation, merge and split. | Mutable glossary record. | Old records remain interpretable; software receives deterministic migration information. |
| Existing-standard bridge | Map proprietary event/API concepts to OpenTelemetry or ISO/industry concepts. | Spreadsheet/manual crosswalk. | Reusable mappings save work without creating competing authoritative concepts. |
| Independent implementation test | Give the published specification and tests to an unaffiliated team. | Founder implementation. | Their server and client interoperate without undocumented knowledge. This is arguably the most important standards proof. |
Proposed proof gates before standardization
These are proposed strategic thresholds, not industry norms.
Before approaching a standards-development organization, the project should be able to show:
Independent implementations: at least three compatible implementations, with at least two maintained outside the originating organization.
External use: several unaffiliated organizations using semantic IDs in real integration paths, not merely demos.
Multiple representation types: production evidence across at least APIs/events and one AI-oriented use case so the generic layer is proven genuinely representation-independent.
Lifecycle evidence: at least one real registry that has undergone deprecation, split/merge, mapping changes and version migration without breaking historical interpretation.
Federation evidence: interoperability among public and private registries operated by independent parties.
Measured economic improvement: pilot data showing that semantic bindings reduce mapping maintenance, migration effort, integration defects or time-to-integrate relative to existing practice.
Resolver evidence: if probabilistic resolution remains part of the project, publish frozen datasets, model versions, precision/recall/abstention curves and hard negatives. Automatic assignment should prioritize precision over coverage because a wrong authoritative semantic bind can be more damaging than an abstention. Embedded Semantics' current methodology already recognizes this need conceptually.
Internationalization evidence: test meaning distinctions, not merely translated aliases. The Embedded Semantics site correctly states that it does not assume every culture partitions meaning identically or that translation is lossless.
Governance evidence: multiple maintainers, public decision logs and actual contested decisions.
Only then does standardization begin to have something real to preserve.
Messaging recommendations
The current public phrase “stable meaning, across languages” is memorable but strategically dangerous because it can be read as claiming that meaning itself is language-independent and universally separable into stable atomic units. The project's own About page is more careful and explicitly disclaims the idea that every culture partitions meaning identically.
The factual positioning should become:
Stable identifiers for governed concepts, with mappings from local representations and multilingual expressions.
For developers:
Attach durable concept identifiers to schemas, events, agent capabilities and memory so references can remain stable when local names or representations change.
For enterprise architects:
A semantic binding layer between local data contracts and governed business concepts.
For standards audiences:
A lightweight registry/federation profile for concept identity and mappings, designed to interoperate with existing vocabularies rather than replace them.
Avoid claims such as:
- universal semantics;
- language-independent meaning;
- ontology replacement;
- “understands meaning”;
- embedding replacement;
- universal AI language;
- global concept authority;
- semantic standard;
- lossless cross-language understanding.
Until external adoption exists, avoid describing the project as an “ecosystem” except aspirationally.
Terminology recommendations
Use “governed concept identifier” or “semantic identifier.” These terms describe the actual primitive without claiming philosophical universality.
Use “semantic foreign key” as an explanatory metaphor. It immediately communicates that an ID points to governed semantic material outside the local schema.
Use “binding” for local-to-concept attachment.
API field → concept
event attribute → concept
agent capability → concept
memory assertion → concept
Use “mapping” for concept-to-concept relationships.
private concept A ↔ public concept B
The distinction will matter enormously in governance.
Use “registry lookup” for deterministic exact authority and “semantic resolution” for an inference process. Do not blur database lookup and model inference.
Use “candidate” for model suggestions. A candidate becomes an accepted binding only after policy/confidence/human rules are satisfied.
Keep “ConceptCode” only as an implementation-friendly compact alias unless it is explicitly bound to a globally scoped URI/IRI.
Avoid “language-neutral.” Prefer “multilingual evidence attached to stable concept identity.”
Avoid using “Embedded Semantics” as the generic industry category. The phrase has prior unrelated use in Semantic Web literature, including earlier work where “embedded semantics” meant semantics embedded into Web content. A distinct generic category such as “governed semantic identifiers” will reduce confusion.
Partnership categories
The first partnerships should maximize independence and interoperability, not logos.
Agent infrastructure: MCP server/client developers, A2A implementers, agent registries and orchestration frameworks. The objective is to prove that semantic IDs enrich rather than compete with current protocols.
Schema/event infrastructure: schema-registry, event-catalog, API-management and integration vendors. Confluent's established schema-evolution model makes this an especially useful adjacent layer for comparison.
Data governance/catalog: DataHub, OpenMetadata and similar communities are natural import/export and binding partners because they already maintain stable glossary entities.
Observability: OpenTelemetry integrations provide a domain with existing semantics and real migration/versioning requirements.
Standards/vocabulary owners: initially as mapping partners—never approach with a proposal to replace their identifiers.
Multilingual/terminology specialists: essential to prevent English-centric concept design and false equivalence.
Academic information-retrieval/NLP teams: useful for independent evaluation of resolver claims and hard-negative benchmarks.
Open governance organizations: useful later for neutral stewardship and IPR processes.
Enterprise design partners: choose firms with significant schema/event/API heterogeneity and enough engineering maturity to measure integration cost before and after adoption.
The project needs independent critics almost as much as partners. A standards-oriented semantic project developed solely among collaborators who already accept its conceptual model will overfit rapidly.
Roadmaps, kill criteria and bibliography
Twelve-month ecosystem roadmap
The current baseline is unusually clear: as of August 23, 2026 the public Embedded Semantics registry contains no published concepts or exact expressions and the semantic runtime is inactive. The first year should therefore be treated as validation from zero, not scale-up.
| Period | Primary objective | Deliverables | Exit criterion |
|---|---|---|---|
| Months 0–3 | Turn the idea into an independently reviewable technical project. | Public specification repository; issue tracker; permissively licensed reference data model; explicit non-goals; canonical URI/IRI identity design; lifecycle RFC; mapping/provenance RFC; security/threat model; initial 100–300 well-scoped concepts in one or two non-regulated test domains; import/export to SKOS-style representation; reproducible registry snapshots. | An external engineer can understand and criticize the architecture without private explanation. |
| Months 4–6 | Prove developer ergonomics. | Reference registry server; CLI; TypeScript/.NET/Python SDKs; schema validator; typed constants; local caching; one MCP semantic-metadata experiment; one event/API binding experiment; DataHub/OpenMetadata importer; resolver benchmark harness. | Two real integrations can add semantic metadata without restructuring their application protocol. |
| Months 7–9 | Prove federation and independence. | Second implementation by an unaffiliated party; namespace delegation; signed registry snapshots; private registry mode; public/private mapping exchange; split/merge/deprecation lifecycle tests; multilingual hard-negative benchmark; public performance results. | Two independent servers and clients interoperate from the written specification and conformance suite. |
| Months 10–12 | Prove economic usefulness and governance. | Several design-partner pilots; before/after integration metrics; external maintainer cohort; public RFC voting/decision process; conformance v1; security audit or structured external security review; first ecosystem report. | At least a few external users choose to continue because the semantic layer reduced a measurable problem, not because they are supporting the research project. |
The twelve-month roadmap should not prioritize a huge registry, formal certification, a new URI scheme, healthcare vocabulary creation or ISO submission.
Three-year ecosystem roadmap
Year one: validate the primitive.
The success question is: Does a semantic foreign-key layer reduce real integration/migration work enough that independent developers voluntarily persist the IDs?
If no, stop trying to create a category.
Year two: build federation and ecosystem economics.
Assuming the first year succeeds, expand connectors and domain profiles rather than the core standard. The core should stabilize while integrations proliferate.
By the end of year two, a healthy ecosystem would have:
- independently operated public and private registries;
- multiple language SDKs;
- schema/API/event/agent/data-catalog integrations;
- a meaningful public mapping commons;
- strong lifecycle tooling;
- neutral technical governance;
- one or more commercial managed-registry offerings, ideally from more than one vendor;
- several external resolver implementations;
- published evidence showing where deterministic identity outperforms LLM-only or embedding-only approaches.
At this stage, the originating company should begin structurally separating open-spec governance from commercial product governance.
Year three: selective standardization.
Only now should the project decide whether there is a generic interoperability layer worth standardizing.
The likely outcome should not be one enormous standard. It should be a small base plus profiles:
Core semantic identity and lifecycle
|
+-- registry federation
+-- mapping/provenance
+-- agent capability binding profile
+-- API/event binding profile
+-- analytics profile
+-- industry-standard mapping profiles
Depending on traction, the generic Web-facing model could move toward W3C; enterprise federation could be suitable for OASIS; agent-specific metadata could be brought to the relevant Linux Foundation/AAIF communities; vertical mappings should remain with existing domain authorities. W3C's insistence on implementation experience and wide review is exactly why this should be a year-three discussion, not a month-three branding goal.
Kill criteria
The project needs explicit evidence that would cause it to stop, shrink or pivot. Otherwise every failure can be rationalized as “the market is not ready.”
The following are proposed go/no-go criteria.
Kill the generic-standard thesis if independent developers do not care about IDs. If pilots repeatedly prefer local aliases, schema metadata or ordinary glossary terms because those are sufficient, the new layer has no reason to exist.
Kill or radically narrow the idea if mapping cost exceeds mapping reuse. If every organization must manually debate and maintain almost every concept mapping, the system merely relocates integration work into a new registry.
Kill the lightweight-layer thesis if most useful concepts require full contextual ontologies. If a stable atomic ID cannot convey enough meaning without an elaborate relationship model, adopt established ontology technologies instead of inventing an inferior version.
Kill the global/public-network thesis if semantics remain overwhelmingly private. Private enterprise registries could still be a useful product, but there may be no public network effect.
Kill automatic semantic resolution if high precision requires impractically high abstention or extensive manual review. The registry can remain useful without an embedding resolver. Do not let a weak AI feature invalidate a sound deterministic registry.
Kill claims of model durability if applications still depend materially on model-specific interpretation after binding. A durable concept ID should actually reduce model coupling.
Kill the standardization effort if a second independent implementation cannot be built from the written specification. Hidden implementation knowledge means there is no standardizable contract yet.
Kill the “unique category” narrative if existing SKOS/glossary/terminology tooling can satisfy the same pilot with comparable developer effort. In that case the better business may be tooling around established standards.
Kill regulated-domain expansion if domain authorities reject the mapping model. Do not enter healthcare, finance or industrial standards by creating parallel semantic authorities.
Kill the ecosystem strategy if governance cannot become genuinely plural. A single-vendor semantic namespace might be a product but should not be represented as open ecosystem infrastructure.
A useful quantitative twelve-month decision framework would be to require, at minimum, evidence of three things:
- Independent demand: multiple unaffiliated teams voluntarily persist semantic IDs in integration artifacts.
- Measured improvement: at least one repeatable workflow shows a material reduction in integration/migration effort or semantic defects versus its existing baseline.
- Independent interoperability: an unaffiliated implementation passes the public conformance suite without founder-specific intervention.
Failure on all three should terminate the generic ecosystem project.
Final strategic judgment
The underlying design principle is sound:
Separate durable semantic identity from mutable representation and probabilistic retrieval.
The public Embedded Semantics documentation articulates that separation well. It keeps ConceptCodes and reviewed registry evidence authoritative, treats vectors as replaceable evidence, explicitly allows abstention and acknowledges that multilingual equivalence is not culturally universal.
But that principle is well precedented. SKOS, UMLS, FHIR, ISO 20022 and other domain standards already demonstrate variations of governed concept identity.
Therefore the strategic question is not:
“Can stable semantic identities work?”
They demonstrably can in governed domains.
The unanswered question is:
Can a lightweight, developer-friendly, federated layer make governed semantic identity useful across ordinary APIs, events, agents, AI memory and enterprise data without imposing the cost of a full ontology or a new centralized semantic authority?
That is the experiment worth running.
The highest-value path is to make Embedded Semantics smaller, less universal and more interoperable:
- reuse URI/IRI identity rather than create a proprietary identifier universe;
- treat ConceptCode as an alias, not the foundation of global identity;
- support SKOS/RDF projection rather than compete with Semantic Web standards;
- map to FHIR, ISO 20022, OPC UA, NIEM, OpenTelemetry and other established authorities rather than duplicate them;
- focus initial product work on agent capabilities, APIs, events, schema integration and durable AI memory;
- make private registries first-class;
- make the public layer federated and mirrorable;
- open the specification, server, test suite and federation protocol;
- monetize operational tooling rather than identifier control;
- separate resolution from authority;
- prove integration savings before seeking a standards venue;
- transfer technical control to plural governance before asking outsiders to make long-lived semantic dependencies.
If those constraints are accepted, there is a plausible new infrastructure layer here.
If the project instead attempts to become a universal semantic language, universal ontology, proprietary global registry or embedding-defined source of truth, existing standards and technologies are substantially better positioned, and the project should be expected to fail.
Bibliography.
| Source | Relevance |
|---|---|
| Embedded Semantics — main architecture | Current concept-identity/vector-authority distinction and exact-resolution positioning. |
| Embedded Semantics — Registry & Runtime Status | Current public implementation status: zero deployed concepts/expressions and inactive semantic runtime. |
| Embedded Semantics — Concept Registry | Confirms current empty public registry. |
| Embedded Semantics — Research & Methodology | Resolver boundaries, hard negatives, abstention and proposed evaluation metrics. |
| Embedded Semantics — About/FAQ | Multilingual/equivalence boundaries and explicit non-claims. |
| JustAnIota — Specification | Registry/schema/validation architecture and authority boundaries. |
| JustAnIota — Registry / Mapping System | Stable-ID, definitions, examples, locale, signing/versioning and registry requirements. |
| JustAnIota — Governance | Current single-publisher governance and future plural-governance boundaries. |
| JustAnIota — Roadmap | Current/planned implementation and evidence gates. |
| Protocol5 — Evidence Workbench / Expression-Concept / UAI distribution | Experimental/implementation context and distinction between evidence and authority. |
| Michael Kappel — AI Memory and Project Portfolio | Related multi-agent, persistent-memory and governance design context. |
| W3C SKOS Reference | Closest open conceptual prior art: URI concepts, multilingual labels, definitions, hierarchy and cross-scheme mappings. |
| W3C RDF 1.2 Concepts | Current IRI-based identity substrate. |
| W3C Process | Wide review, implementation experience, consensus and Registry Track requirements. |
| NLM UMLS — Unique Identifiers | Direct precedent for distinguishing meanings, strings and concept IDs. |
| HL7 FHIR — CodeSystem | Canonical code-system identity, versioning, concept permanence and terminology publishing. |
| ISO 20022 Data Dictionary and governance | Governed, uniquely identified business concepts and formal maintenance procedures. |
| OPC Foundation — Companion Specifications | Industrial semantic interoperability and cross-vendor information models. |
| OPC UA Robotics | Existing robotics semantic-information-model precedent. |
| OpenTelemetry Semantic Conventions | Established semantic naming, meaning, stability and governance in observability. |
| NIEM Open | Government semantic/syntactic interoperability and multi-representation model. |
| OASIS / NIEM ISO submission | Contemporary example of ecosystem work maturing toward international standardization. |
| Model Context Protocol, July 2026 specification | Current tool/data interoperability substrate for AI systems and agent-adjacent integration. |
| A2A / Linux Foundation | Open agent-to-agent interoperability and neutral governance precedent. |
| Agentic AI Foundation | Current neutral-foundation governance for major agent infrastructure projects. |
| Confluent Schema Registry | Structural contract and schema-evolution competitor/complement. |
| DataHub glossary model | Governed business terms, stable identifiers and attachment to enterprise assets. |
| OpenMetadata glossary model | Existing enterprise controlled-vocabulary and semantic-graph capability. |
| Palantir Foundry Ontology | Broader enterprise ontology/platform alternative. |
| ISO standards-development process | Multi-stakeholder, consensus-based standardization model relevant to eventual maturity. |
| OASIS Open Project rules | Possible neutral governance model for code-plus-specification development. |