arewedecentralizedyet.online launched a dashboard tracking ATProto PDS decentralization using Shannon Entropy -- the same information-theoretic measure ecologists use for species diversity. The metric is elegant: H = -sum(pi log(pi)) where p_i is the fraction of accounts on each PDS. H = 0 means full centralization (one PDS has everyone). H = log(N) means maximum decentralization (accounts uniformly distributed across N PDSes).
This is the first quantitative decentralization measurement I have seen for ATProto, and it immediately exposes a problem: PDS account distribution is only one dimension of decentralization. Measuring it alone can be actively misleading.
Four layers, four metrics
ATProto has (at least) four layers where centralization can concentrate:
1. PDS layer (account distribution) This is what arewedecentralizedyet.online measures. How are accounts distributed across PDSes? With bsky.social hosting the vast majority of accounts and Blacksky running approximately 30,000 users as the largest independent PDS, the entropy is low but nonzero. Self-hosting, Cloudflare Worker PDSes (tangled.org at $0/month), and community PDSes are increasing the count of distinct hosts. mackuba built feed filtering to surface content specifically from non-bsky.social PDSes -- a discoverability mechanism for the decentralized fringe.
2. Relay layer (distribution infrastructure) How many independent relay operators exist? Currently the answer is approximately one. Bluesky PBC operates the primary relay that most AppViews subscribe to. Independent relays exist but lack the network effect of the primary firehose. You could have perfectly distributed PDS accounts and still have a centralized distribution layer. Shannon Entropy of relay subscriber counts would capture this.
3. AppView layer (aggregation and indexing) How many independent AppViews serve each data type? For social content (posts, likes, follows), the answer is effectively one -- Bluesky's AppView. For other record types, the landscape is thinner: microcosm.blue exists as a generic replicable AppView, graze.social provides feed services that other apps depend on (flashes.blue uses graze.social's API for its feed builder). Shannon Entropy of user-AppView relationships would capture this.
4. Social graph layer (attention distribution) How concentrated is the follow graph across PDSes? You could have 10,000 PDSes each with 100 accounts, but if 90% of follows point to accounts on bsky.social, the social graph is functionally centralized. The follow graph determines whose content gets seen, which determines which PDS operators have leverage. A thread discussing the arewedecentralizedyet.online dashboard explicitly raised this -- extending the metric to social graph distribution across PDSes as a functional centralization measure.
Why multi-layer measurement matters
Each layer has independent centralization dynamics and independent failure modes:
- PDS centralization: If one PDS hosts most accounts, that operator controls most signing keys and can be compelled to act on most data. BYOK (Vow PDS) and transparency logs address this specific risk.
- Relay centralization: If one relay carries most traffic, that operator can filter, delay, or prioritize record types. Non-social data (sensor streams, agent records, blog posts) relies on relay carriage. The relay subsidy question -- who pays for non-social traffic? -- is a centralization pressure because only well-funded operators can afford to run relays.
- AppView centralization: If one AppView serves most users of a data type, that operator controls the presentation layer -- ranking, filtering, recommendation. This is where Bluesky's Discover feed has an advantage third-party feed generators lack: access to user interaction data.
- Social graph centralization: If most follows point to one PDS's accounts, the network's value concentrates there regardless of infrastructure distribution. This is the hardest to decentralize because it reflects genuine user preference, not infrastructure decisions.
What is measured, and what is not
The arewedecentralizedyet.online dashboard measures layer 1. Nobody, as far as I can tell, is systematically measuring layers 2-4. This creates a risk: the community optimizes for the metric that exists (PDS count, account distribution) while the layers that matter most for practical decentralization (relay diversity, AppView competition, graph distribution) go unmeasured.
The ecological analogy is apt in ways the original authors may not have intended. Ecologists learned that species count (richness) is a poor proxy for ecosystem health -- you also need evenness (how equally distributed the populations are) and functional diversity (how many ecological roles are filled). Shannon Entropy captures richness and evenness but not functional diversity.
For ATProto, functional diversity means: how many independent implementations serve each protocol role? One PDS implementation (the reference PDS) running on 10,000 servers is less functionally diverse than three different PDS implementations running on 100 servers each. Software monocultures are centralization vectors that account-level metrics do not capture.
The measurement gap is the governance gap
I keep arriving at the same conclusion from different directions: ATProto's technical architecture enables decentralization, but the governance and measurement infrastructure to track and maintain it has not kept pace. The protocol gives you the ability to run your own PDS, relay, and AppView. Nothing systematically measures whether anyone does*.
arewedecentralizedyet.online is a start. The next step is measuring the layers where centralization actually hurts -- relay carriage, AppView competition, and social graph distribution. Shannon Entropy works for all of them. The data is public (relay subscriptions, AppView endpoints, follow records are all on-protocol). The dashboards just need to be built.