Notes for estimating an atproto CAR file size

By Archived fry69 (@fry69.altq.net)
Published:

• Findings

How to Get an Estimate

Example (using Kysely inside the PDS codebase):

const totalBlockBytes = await actorDb.db
    .selectFrom('repo_block')
    .select(({ fn }) => fn.sum('size').as('sum'))
    .if(since, (qb, s) => qb.where('repoRev', '>', s))
    .executeTakeFirst()
    .then((row) => Number(row?.sum ?? 0))

  const blockCount = await actorDb.db
    .selectFrom('repo_block')
    .select(countAll.as('count'))
    .if(since, (qb, s) => qb.where('repoRev', '>', s))
    .executeTakeFirst()
    .then((row) => Number(row?.count ?? 0))

  const estimatedCarBytes =
    headerBytes +
    blockCount * (36 + averageVarintLen) + // per-block overhead
    totalBlockBytes

Because SQLite pages, indexes, and WAL files add their own overhead, the on-disk .sqlite file will always be larger than the exported CAR. The only reliable way to relate them is through the summed repoblock.size described above. If this needs to be exposed publicly, you’d add an endpoint (or extend getRepoStatus) that runs those aggregate queries and returns the estimate before streaming.