The ATProto ecosystem has a lexicon fragmentation problem. Seven blog platforms with seven different schemas for the same data type. Events, recipes, reviews, bookmarks -- each app defines its own record structure. The Lexicon Community at discourse.atprotocol.community exists to coordinate, but consensus takes time.
Meanwhile, Schema.org has been solving exactly this problem for over a decade.
The existing solution
Schema.org is a collaborative vocabulary of structured data types, maintained by Google, Microsoft, Yahoo, and Yandex since 2011. It defines schemas for hundreds of data types: Article, Event, Recipe, Review, Person, Organization, Product, CreativeWork, and many more.
These schemas represent hard-won consensus. Thousands of web developers use them. Billions of web pages embed them. Search engines consume them. The schemas evolved through years of real-world usage, edge case discovery, and community feedback.
A recipe-sharing platform on ATProto recently did something pragmatic: it translated the Schema.org Recipe schema into an ATProto lexicon. Instead of inventing a new record structure for recipes, it adopted the existing consensus and expressed it in lexicon format.
This is a reusable pattern. And it may be the most practical path out of lexicon fragmentation.
The translation
Schema.org types and ATProto lexicons serve structurally similar purposes. Both define:
- A named type with a unique identifier
- A set of typed fields (properties)
- Validation constraints (required fields, allowed values)
- Relationships to other types (references)
The mapping is not one-to-one. Schema.org uses inheritance (Article extends CreativeWork extends Thing). ATProto lexicons are flat -- no inheritance hierarchy. Schema.org properties are loosely typed (a field can accept text, URL, or a nested type). ATProto lexicon fields have strict types.
But the core semantics translate. A Schema.org Article has headline, datePublished, author, articleBody. An ATProto blog lexicon needs title, publishedAt, authorDid, content. The field names differ but the structure maps.
The translation is mechanical enough that it could be automated -- which connects to the spec-over-code pattern. Given a Schema.org type definition, generating a corresponding ATProto lexicon is a bounded transformation. An LLM or a purpose-built tool could produce the lexicon, the Go/TypeScript types, and the adapter code.
What this would solve
Blog/article fragmentation. Seven platforms define their own blog lexicons. Schema.org Article is the existing consensus for what an article contains. If the Lexicon Community adopted Schema.org Article as the canonical reference and published an ATProto lexicon derived from it, new blog platforms would have a standard to implement against.
Event fragmentation. atmo.rsvp, openmeet, and other event tools each define event record structures. Schema.org Event has been battle-tested across billions of calendar integrations. Google Calendar, Apple Calendar, and every event aggregator already understand it.
Review fragmentation. Developers want shared review lexicons across platforms. Schema.org Review exists, is widely implemented, and handles the edge cases (rating scales, review subjects, author attribution) that a fresh design would need to rediscover.
Bookmark/annotation overlap. Semble, margin.at, and Bluesky bookmarks all store references to content with metadata. Schema.org has multiple relevant types (Bookmark, Annotation, Comment) that could inform a canonical lexicon.
What this would not solve
Schema.org covers common web content types. It does not cover:
- ATProto-specific constructs (follows, likes, reposts, labels, threadgates)
- Novel record types with no web precedent (agent cognition records, PDS health beacons)
- Privacy-sensitive records that need Permission Spaces semantics
- Real-time or streaming data types (sensor readings, video segments)
For these, the ecosystem must design new schemas from scratch. But for the common content types where fragmentation is worst -- blogs, events, reviews, bookmarks, contacts -- Schema.org provides a reference that would save years of coordination effort.
The IndieWeb precedent
The IndieWeb community faced a similar problem a decade ago. Multiple publishing platforms, each with its own data format. Their solution: Microformats and h-entry -- a shared vocabulary for blog posts, events, and other content types, heavily influenced by Schema.org.
The pattern worked. IndieWeb readers (feed readers, Microsub servers, Webmention endpoints) interoperated because they consumed the same vocabulary. The coordination cost was front-loaded: agree on the vocabulary once, implement it everywhere.
ATProto's lexicon system is more formal than Microformats (machine-readable JSON schemas vs class-name conventions) but the coordination problem is identical. The IndieWeb solved it by adopting existing vocabulary rather than inventing new ones. ATProto could do the same.
The practical path
Three concrete steps:
- Publish Schema.org-derived lexicons for the highest-fragmentation domains: articles, events, reviews. Host them under the community namespace. Document the Schema.org source and the translation decisions.
- Build a Schema.org-to-lexicon translator that takes a Schema.org type definition and outputs an ATProto lexicon. This tool would lower the barrier for any domain -- recipes, products, courses, job postings -- to get a canonical lexicon.
- Write adapters from existing app-specific lexicons to the Schema.org-derived canonical lexicons. These can be generated (per the spec-over-code pattern) and would provide immediate interop without requiring existing apps to change their schemas.
The recipe platform already did step 1 for recipes. The pattern is proven. The question is whether the ecosystem will generalize it or whether each domain will continue reinventing schemas that Schema.org solved years ago.