Blob: docs/streaming-push.md
Streaming Push and R2-First Repository Design
Status
Streaming is production. Legacy paths removed in closure release. All repos are streaming. The storage-mode concept, rollback backfill, unpack, hydration, and legacy receive code have been removed. See MIGRATION-STREAMING-PUSH.md for upgrade guidance.
This document specifies the streaming push design that replaced the unpack-and-hydrate pipeline.
The new design makes R2 .pack and .idx objects the only required source of truth for Git object data. The repository Durable Object remains the source of truth for metadata only: refs, HEAD, pack catalog state, and receive/compaction leases.
The design is intentionally conservative about complexity:
- No Git protocol changes.
- No repo metadata stored in R2.
- No exact
pack_objects(pack_key, oid)correctness dependency. - No correctness dependency on DO
obj:*or R2 loose mirrors. - No new distributed transaction.
- No new tunable env vars beyond a single Queue binding.
- No object-level garbage collection in the current design.
Why This Exists
The current implementation buffers the entire receive-pack request body in memory, indexes through isomorphic-git over an in-memory filesystem, then unpacks to loose objects in the Durable Object and mirrors them to R2. That model does not scale and violates the intended ownership boundary:
- the DO is doing data-plane work instead of metadata-plane work;
- correctness depends on loose object materialization;
- fetch/object reads still depend on loose-object semantics;
- exact pack membership is stored in SQLite even though R2
.idxalready contains the authoritative mapping; - background hydration is a second correctness mechanism layered on top of unpacking.
The new model replaces all of that with immutable packs in R2, worker-local range reads, and compact metadata in the DO.
Goals
- Streaming-first receive path with bounded memory.
- R2-first data model: packs in R2 are authoritative for object data.
- DO metadata-only correctness boundary.
- No externally visible Git protocol changes.
- Existing packed repos continue to work.
- Existing repos can lose all loose objects and still remain correct.
- Fetch becomes naturally pack-first and streaming.
- Heavy lifting moves out of the DO and into Workers/Queue consumers.
- The closure release leaves a single streaming-only end state.
Non-Goals and Deliberate Simplifications
These are intentional, not omissions:
- The design does not implement full Git garbage collection of unreachable objects. Compaction preserves the union of source-pack objects. This matches current retention behavior more closely and avoids a repo-wide mark phase.
- The fetch/read path loads whole
.idxfiles into memory when needed. This is acceptable because the representative.idxis about 3 MB while the representative.packis about 42 MB. The constraint is “do not buffer the pack”, not “never buffer an idx”. - Push concurrency is simplified to one active receive lease per repo. The current one-active plus one-queued unpack model is removed because unpacking is removed. Returning
503 Retry-After: 10to concurrent pushes is Git-compatible and materially simpler. - UI raw/blob reads may buffer an individual packed object when it is delta-resolved and not already present in the optional loose cache. This is acceptable because the streaming-first requirement is about pack ingress/egress paths, not ad hoc UI reads; the UI already enforces size caps.
- Bloom filters are not introduced initially. Active pack count is bounded by compaction, and the
.idxfanout table already gives an exact, compact lookup structure.
Source of Truth
Data plane
Required for correctness:
do/<do-id>/objects/pack/*.packdo/<do-id>/objects/pack/*.idx
Optional caches only:
- DO
obj:* - R2 loose mirrors under
do/<do-id>/objects/loose/*
Metadata plane
Required for correctness:
- DO KV:
refs,head,refsVersion,packsetVersion,nextPackSeq, receive/compaction leases - DO SQLite: active/superseded pack catalog
Legacy data removed in the closure release (stale keys may remain on disk but are not read):
pack_objects,hydr_cover,hydr_pending(SQLite tables dropped)lastPackOids,lastPackKey,packList,unpackWork,unpackNext(DO KV keys)obj:*(DO loose object cache)
Key Design Decisions
1. .idx is treated as data-plane state, not repo metadata
Refs, HEAD, pack catalog, and leases remain in the DO. Standard Git .idx files live next to .pack files in R2 because they are required to address immutable pack contents efficiently. No JSON manifests or repo catalog records are stored in R2.
2. Receive writes pack data first, commits metadata last
The Worker only publishes a new pack to the repo catalog after:
- the
.packobject is fully written to R2; - the
.idxis successfully built and written to R2; - pack integrity and thin-pack base validation pass;
- ref/update validation passes.
This avoids a distributed transaction. R2 objects are immutable blobs; the DO metadata only points to fully written pack pairs.
3. Fetch and object reads stop depending on loose objects
All fetch planning, object existence checks, commit/tree/tag/blob reads, raw/blob views, diff, and merge-history traversal must use a shared pack-first object store in Worker code. DO object RPCs become compatibility shims or are removed from callers.
4. Hydration is replaced by queue-driven pack compaction
The system no longer unpacks to loose objects or builds “hydration packs”. It keeps a bounded set of immutable active packs and periodically compacts older packs into larger packs in a background Queue consumer.
5. Compaction is tiered, not full-repo repack
The initial compaction strategy is LSM-like:
- receive packs enter tier 0;
- when a tier exceeds the fixed fan-in, the oldest packs in that tier are merged into one pack in the next tier;
- active pack count remains bounded;
- the algorithm does not require a full-repo rewrite on each push.
This is simpler than full snapshot compaction and good enough for the target workload.
Assumptions
- The representative push remains within the Worker request body limit. The supplied sample is about 42 MB, which is under the documented minimum 100 MB request-body limit on paid Cloudflare plans.
- Workers and Durable Objects run with a 128 MB memory limit; request/response streaming is available; R2
put()accepts aReadableStream; R2 ranged reads are strongly supported in Workers; Queue consumers have 15 minutes of wall time. - Active pack count is kept low by compaction, so loading a handful of
.idxfiles into memory is safe and simpler than inventing a second metadata structure. - Existing repos that already had packs in R2 were the primary cutover migration case. Closure assumes that any loose-only repos already completed the required one-time pack migration before this release.
Cloudflare Constraints Used By The Design
These are the platform facts the design relies on:
- Workers and Durable Objects have a 128 MB memory limit.
- HTTP request handling has no hard wall-time limit while the client stays connected.
- Queue consumers have a 15 minute wall-time limit.
- Workers Paid supports up to 10,000 subrequests per invocation and up to six simultaneous outgoing connections.
- R2
put()is strongly consistent and accepts aReadableStream. - R2
get()supports ranged reads byoffsetandlength. - R2 multipart uploads exist, but require 5 MiB parts except the last part.
The receive path prefers a single Worker-to-R2 put() when the pack length is known ahead of time. When the incoming Git client uses chunked transfer and the pack length is not knowable up front, the receive path falls back to Worker-to-R2 multipart upload with fixed-size parts.
New Metadata Model
New SQLite table: pack_catalog
Add a new table via src/do/repo/db/schema.ts and DAL helpers in src/do/repo/db/dal.ts.
Required columns:
packKey TEXT PRIMARY KEYkind TEXT NOT NULLreceivecompactlegacy
state TEXT NOT NULLactivesuperseded
tier INTEGER NOT NULLseqLo INTEGER NOT NULLseqHi INTEGER NOT NULLobjectCount INTEGER NOT NULLpackBytes INTEGER NOT NULLidxBytes INTEGER NOT NULLcreatedAt INTEGER NOT NULLsupersededBy TEXT NULL
Indexes:
- active packs ordered by
state,seqHi DESC - active packs in a tier ordered by
state,tier,seqLo
Notes:
seqLoandseqHiare catalog ordering markers, not Git semantics.- Active packs must always represent a disjoint cover of receive-sequence ranges.
- New receive packs start with
seqLo = seqHi = nextPackSeq. - A compacted pack gets
seqLo = min(source.seqLo)andseqHi = max(source.seqHi).
DO KV keys that remain authoritative
refsheadrefsVersionpacksetVersionnextPackSeqreceiveLeasecompactLeasecompactionWantedAt
Lease contents:
receiveLease = { token, createdAt, expiresAt }compactLease = { token, createdAt, expiresAt }
Fixed code constants, not env vars:
- receive lease TTL: 30 minutes
- compaction lease TTL: 20 minutes
Expired leases are cleared:
- on the next
beginReceive()/beginCompaction()call; and - by a lightweight DO alarm cleanup path.
Legacy compatibility mirrors (removed in closure)
packList, lastPackKey, and lastPackOids were maintained as mirrors during rollout for rollback compatibility. They have been removed in the closure release. pack_catalog is the sole authority for pack metadata.
Cross-System Mutation Rules
Rule 1
Only the DO mutates repo metadata.
Rule 2
Workers and Queue consumers may write immutable R2 blobs before metadata commit, but those blobs are not visible to the repo until the DO commits the catalog row.
Rule 3
Deleting unreferenced staged pack blobs and superseded R2 pack pairs is best-effort and retryable. Failure to delete must never affect correctness.
Rule 4
Compaction never mutates source packs in place. It writes a new pack, then atomically swaps catalog rows in the DO.
Rule 5
Fetch and UI reads operate on a read-only snapshot of the active catalog. They never mutate repo state.
Rule 6
Queue delivery is a hint, not the durable record of pending compaction. The durable record is compactionWantedAt in DO metadata. Queue delivery failure may delay compaction, but it must not lose the need for compaction.
Receive Path
Overview
The Worker handles receive-pack end to end. The DO is used only for:
- acquiring a receive lease;
- reading current refs/HEAD/version metadata;
- committing ref changes and the new pack row;
- clearing the lease.
The DO does not expose any HTTP receive endpoint. All receive coordination uses typed RPCs.
Client-visible behavior
Unchanged:
- Git Smart HTTP request/response format.
report-status.- delete-only pushes.
- stale
old-oidrejection. - invalid ref rejection.
- thin-pack acceptance when bases exist.
Changed internally:
- concurrent push handling becomes “one active receive lease per repo”; additional pushes get
503 Retry-After: 10. X-Repo-ChangedandX-Repo-Emptyare computed in the Worker after DO finalization, not forwarded from a DO HTTP endpoint.
Step-by-step algorithm
- Worker calls
stub.beginReceive()before reading the body. - If a valid receive lease already exists, the Worker immediately returns
503 Retry-After: 10. - If the lease is granted, the Worker incrementally parses the pkt-line command section from
request.bodyuntil the flush packet. - The remaining bytes are treated as the raw pack stream.
- The Worker writes the raw pack stream directly to a staged R2 key such as
do/<id>/objects/pack/pack-rx-<leaseToken>.pack. - During the upload stream, the Worker validates only stream-level invariants that are cheap to check in one pass:
- pack starts with
PACK - version is 2
- trailer SHA-1 matches the streamed body
- byte count matches the actual received length
- pack starts with
- After the R2 upload resolves, the Worker runs the new indexer against the staged pack in R2 and writes
pack-rx-<leaseToken>.idx. - The Worker runs thin-pack base validation and ref-target connectivity validation using the new pack-first object store and the active pack catalog snapshot.
- The Worker calls
stub.finalizeReceive(...). - The DO atomically:
- rechecks the lease token;
- rechecks
old-oidexpectations against current refs; - allocates
nextPackSeq; - inserts the new active pack row;
- updates
refs,head,refsVersion, andpacksetVersion; - sets or refreshes
compactionWantedAtif the catalog now violates the compaction policy; - clears the receive lease;
- returns whether compaction should be queued.
- The Worker returns the pkt-line
report-statusresponse. - If compaction should run, the Worker enqueues one idempotent Queue message
{ repoId }inctx.waitUntil(...).
Failure handling
If any step before finalizeReceive() fails:
- the Worker calls
stub.abortReceive(leaseToken)best-effort; - the staged
.packand.idxare deleted best-effort; - the response is either
400,415,500, or503depending on failure class.
If finalizeReceive() fails stale-old-oid validation:
- the Worker deletes staged
.packand.idxbest-effort; - the Worker returns a normal Git
report-statusrejection, not an HTTP 409.
If the client disconnects after finalizeReceive() succeeds:
- refs and catalog remain committed;
- a retry from the client is expected to fail stale-old-oid, which is standard Git behavior.
New Pack Indexer
Why a new indexer exists
The current isomorphic-git path requires the whole pack in memory. The replacement must index from R2 without buffering the entire pack.
Design
The new indexer lives in src/git/pack/indexer/ and runs entirely in Worker code.
The indexer is two-stage:
scanPack()resolveDeltasAndWriteIdx()
Implementation note:
scanPack()cannot rely onDecompressionStreamalone because Git pack entries are concatenated deflate streams and the scanner must know exactly how many compressed bytes each entry consumed;- the implementation therefore needs a byte-accounting inflate cursor in JS/WASM that exposes end-of-stream position for each packed entry.
scanPack()
Input:
- staged R2
.pack - active pack catalog snapshot for external-base lookup
Behavior:
- reads the pack sequentially from R2 in fixed-size ranges;
- validates header and trailer;
- records, for each object:
- ordinal
- offset
- packed type
- header length
- packed span end
- CRC32 of the raw packed entry
- base offset for
OFS_DELTA - base oid for
REF_DELTA - result size for delta objects
- computes final oid immediately for non-delta objects by inflating the object payload stream and hashing
"<type> <size>\\0<payload>".
Output:
- a compact object-entry table;
offset -> indexlookup;- the list of non-delta resolved objects;
- the list of unresolved delta objects.
Implementation constraint:
- The object-entry table must use typed arrays or chunked binary buffers, not arrays of large JS objects.
- OIDs are stored as raw 20-byte values internally; hex strings are materialized lazily.
resolveDeltasAndWriteIdx()
Behavior:
- resolves unresolved delta objects in dependency order;
- supports:
- in-pack
OFS_DELTA - in-pack
REF_DELTA - thin-pack external
REF_DELTAbases
- in-pack
- rejects any delta whose base cannot be resolved from either:
- an already-scanned in-pack object, or
- the active pack catalog snapshot.
Memory strategy:
- resolved base payloads are kept in an LRU cache with a hard byte budget;
- evicted base payloads may be recomputed from the pack/object store if needed later;
- correctness does not depend on cache hits.
Delta application:
- parse the delta header varints;
- verify the declared base size;
- materialize the result payload into a pre-sized buffer;
- compute the final oid from
"<baseType> <resultSize>\\0<resultPayload>".
After all object oids are resolved:
- sort entries by oid;
- emit a standard Git idx v2 file;
- write it to R2;
- return
objectCount,idxBytes, and the parsedIdxView.
Why this is feasible
For the representative pack:
- pack: about 42 MB
- idx: about 2.7 MB
- objects: about 97k
The design never holds the whole pack in memory. It only holds:
- the compact entry table;
- a bounded payload cache for delta bases;
- the final idx buffer.
That fits comfortably within the 128 MB memory limit for the target workload.
Connectivity Validation
Validation remains intentionally limited to what the current product already enforces.
Pack-level validation
Reject the push if any packed object is structurally invalid or any thin-pack external base cannot be resolved from the active catalog snapshot.
Ref command validation
Keep the existing rules:
- reject
HEADupdates; - reject invalid ref names;
- delete requires existing ref and matching
old-oid; - create requires zero
old-oid; - update requires matching
old-oid; - no partial apply: all commands succeed or none apply.
Target-object connectivity validation
For each non-delete newOid:
- resolve tags transitively, with a hard depth limit of 8;
- if the final type is
commit, require:- object exists;
- root tree exists;
- each parent exists;
- if the final type is
treeorblob, require it exists.
The design does not add full fsck or full tree-walk validation.
Fetch and Read Path
New shared object store
Create src/git/object-store/ and move all correctness-critical object reads to it.
Required APIs:
loadActivePackCatalog(repoId)loadIdxView(packKey)findObject(oid)hasObjectsBatch(oids)readObject(oid)readObjectRefsBatch(oids)readBlobStream(oid)iterPackOids(packKey)
Catalog loading
getPackCandidates() stops using pack_objects and R2 listing as correctness sources.
New behavior:
- ask the DO only for the active pack catalog;
- cache the result in request memo;
- sort active packs by
seqHi DESC, then bytier DESC.
R2 listing remains migration-only fallback, never the primary read path.
Idx loading
IdxView replaces pack_objects as the membership source.
Behavior:
- load the entire
.idxobject into memory; - keep raw fanout, raw name table, and raw offset tables;
- build
oid -> indexlazily or via binary search over the raw name table; - build
offsetToIndexandnextOffsetonce per loaded idx.
Object resolution
findObject(oid):
- iterate active packs in catalog order;
- use the loaded
.idxto check membership; - return the first hit as
(packKey, objectIndex, offset, nextOffset).
readObject(oid):
- locate the object by
.idx; - read the object entry header and compressed payload span from the pack via range reads;
- if base object type, inflate and return payload;
- if delta, recursively resolve base and apply delta;
- memoize per request.
Closure planning
Replace all current dependencies on DO hasLooseBatch(), getPackOids*(), and getObjectRefsBatch().
New rules:
findCommonHaves()useshasObjectsBatch()over the active pack catalog.buildUnionNeededForKeys()unions.idxmembership from the selected packs.computeNeededFast()uses worker-localreadObjectRefsBatch()over packed objects.
Streaming fetch
The existing streaming assembler remains the fetch data path, with two changes:
- pack discovery comes from the active pack catalog rather than
packList + pack_objects; - all object/membership lookups come from the worker-local object store.
This means:
- single-pack fetch remains streaming;
- multi-pack fetch remains streaming;
- fetch becomes correct even when all loose objects are removed.
UI and admin read paths
The following routes must stop depending on DO object RPCs for correctness:
- tree
- blob
- raw
- rawpath
- commits
- commit diff
- merge fragment expansion
Required behavior:
- if a loose cache entry exists, it may be used;
- if not, the route must read from packs;
- if a blob is delta-resolved and over the existing UI size cap, preserve the current “too large” behavior.
readBlobStream()must preserve correctness for raw/blob routes, but it is not required to be zero-copy; it may materialize an individual packed object or delta base chain as long as it never requires whole-pack buffering.
Optional Loose Cache
The DO may continue to cache loose objects, but only as an opportunistic cache.
Rules:
- no fetch or receive correctness may depend on it;
- compaction does not require it;
- deleting all loose data must not break clones, fetches, or UI reads.
Permitted uses:
- hot tree/commit/tag cache
- hot raw/blob cache
- debug-only loose fallbacks
Background Compaction
Why compaction exists
Removing unpack/hydration means packs in R2 must themselves remain a manageable serving set. Compaction is the replacement for “loose objects + hydration packs”.
Queue model
Add one Queue binding, for example REPO_MAINT_QUEUE.
Producer:
- the main Worker after successful receive finalization;
- admin compaction trigger;
- migration jobs.
Consumer:
- the same Worker export, with
queue()implemented insrc/git/compaction/run.ts.
Recommended Queue config:
max_batch_size = 1max_batch_timeout = 1- default retries are acceptable
No new env vars are required.
The DO alarm remains in use, but only for lightweight metadata work:
- expire stale receive and compaction leases;
- retry queue re-arm when
compactionWantedAtis set and no compaction lease is active; - never perform pack indexing, unpacking, or compaction itself.
Compaction policy
Fixed constants in code:
- fan-in: 4
- one compaction lease per repo
Policy:
- tier 0 contains receive packs;
- if more than four active packs exist in a tier, compact the oldest four in that tier;
- output one pack in the next tier;
- source packs become
supersededonly after DO commit. - no compaction lease may be granted while a receive lease is active for the same repo.
- receive has priority over compaction; if a receive lease becomes active after compaction starts,
commitCompaction()must fail and the queue worker must retry later.
Compaction algorithm
- Queue consumer calls
stub.beginCompaction(). - The DO either returns “no work” or returns:
- lease token
packsetVersion- selected source packs
- full active catalog snapshot
- target tier
- The worker computes
needed = union(source pack idx membership). - The worker calls the existing or updated streaming assembler with:
packKeys = full active catalogneeded = union(source pack membership)
The full active catalog is passed so that delta bases outside the source set are pulled in automatically. The output pack therefore becomes self-contained enough to replace the source packs.
- The output stream is written directly to a staged compacted-pack key in R2.
- The new indexer runs against the staged compacted pack and writes the staged
.idx. - The worker calls
stub.commitCompaction(...). - The DO atomically:
- rechecks the lease token;
- rechecks
packsetVersion; - rechecks that the selected source packs are still active and unchanged;
- marks the new pack active;
- marks source packs superseded;
- updates
packsetVersion; - clears or refreshes
compactionWantedAtbased on the post-commit catalog state; - clears the compaction lease.
- The worker deletes superseded R2 blobs best-effort in
waitUntil(...).
What compaction does not do
It does not:
- rewrite refs;
- delete unreachable objects from within a pack;
- require loose objects;
- require hydration state or
pack_objects.
Existing Repositories
Existing packed repos (pre-closure migration, now completed)
Migration was automatic during the cutover phase:
- On DO startup or first access, if
pack_catalogwas empty:- seeded from the union of
lastPackKey,packList, and the full R2.packlisting under the repo prefix (these keys no longer exist post-closure); - ignored
pack_objectsfor correctness.
- seeded from the union of
- For each discovered pack:
- required
.packand.idxto exist; - parsed
.idxfanout to get object count; - inserted a
legacyactiverow intopack_catalog; - assigned synthetic sequential
seqLo = seqHivalues.
- required
- Legacy
pack-hydr-*packs were inserted as normal active packs.
Existing loose-only repos (pre-closure migration, now completed)
These could not switch directly to repoStorageMode = streaming during the cutover migration.
Required migration:
- Traverse reachable objects from current refs using the legacy loose path.
- Build one pack in R2 with a streaming non-delta pack writer.
- Index it with the new indexer.
- Insert it into
pack_catalogas alegacyactive pack. - Leave loose data untouched.
Only after that migration completed could the repo be promoted during cutover.
Emergency rollback (removed in closure release)
The emergency backfill tool described here was used during the phase-3 canary rollout and has been removed in the closure release. See docs/streaming-push-closure-plan.md for details.
Compatibility Surface
Routes and headers
Unchanged public Git routes:
GET /:owner/:repo/info/refsPOST /:owner/:repo/git-upload-packPOST /:owner/:repo/git-receive-pack
Internal headers:
X-Repo-ChangedX-Repo-Empty
These remain worker-internal and continue to exist if callers still use them.
RPC compatibility (historical)
The following RPCs were removed across phases 1-5 and the closure release:
getObjectStream,getObject,getObjectSize— replaced by worker-local pack-first readshasLooseBatch— replaced byhasObjectsBatchover pack cataloggetObjectRefsBatch— replaced by worker-local packed-object batch readergetPackLatest,getPackOids,getPackOidsBatch— replaced bypack_catalogas sole authoritygetUnpackProgress— replaced bybeginReceive()lease acquisitiongetRepoStorageMode,setRepoStorageModeGuarded— storage-mode concept removed
Progress and activity UI
The current UI renders unpack progress in multiple places. That caller graph must be migrated explicitly.
Replacement contract:
- replace
getUnpackProgress()insrc/common/progress.tswithgetRepoActivity(); getRepoActivity()returns a neutral idle state plus two optional live states:receivingcompacting
- the existing progress banner remains in place, but its text changes:
receiving: “Receiving push…”compacting: “Compacting packs…”- idle: render nothing.
Files that must migrate together:
src/common/progress.tssrc/routes/ui/overview.tssrc/routes/ui/tree.tssrc/routes/ui/commits.tssrc/routes/ui/adminPage.tssrc/client/components/ProgressBanner.tsxsrc/client/pages/OverviewPage.tsxsrc/client/pages/TreePage.tsxsrc/client/pages/CommitsPage.tsxsrc/client/pages/AdminPage.tsx
During phase 2, these files may continue to render a banner, but they must stop interpreting unpackWork, queuedCount, or hydration presence as correctness signals.
Admin endpoints
POST /:owner/:repo/admin/compact— returns a compaction plan by default and triggers compaction whendryRun === falseDELETE /:owner/:repo/admin/compact— clears queued compaction work for the repo
The /admin/hydrate aliases and /admin/storage-mode endpoints were removed in the closure release.
Current pack-deletion admin behavior must also change:
DELETE /:owner/:repo/admin/pack/:packKey- may only delete
supersededpacks; - must reject deletion of
activepacks by default; - any forced delete of an
activepack remains an explicit admin-only break-glass path and is outside normal operating assumptions.
- may only delete
Debug/admin fields
Replace loose/unpack/hydration-centric fields with:
receiveLeasecompactionactivePackssupersededPackspackCatalogVersion
Legacy debug fields (unpacking, queuedCount, lastPackOids, repoStorageMode, hydration*) were removed in the closure release.
Current debug endpoints must stay functional, but they must read through the pack-first object store:
GET /:owner/:repo/admin/debug-stateGET /:owner/:repo/admin/debug-commit/:commitGET /:owner/:repo/admin/debug-oid/:oid
Admin UI migration (completed):
- The admin island uses compaction-derived status exclusively
StorageModeCardandHydrationCardcomponents have been removed- Pack-removal warnings reference compaction supersession, not hydration
What Was Removed (Closure Release)
The following subsystems were removed in the closure release:
- DO
POST /receiveendpoint - Background unpacking (alarm-driven loose object extraction)
- Hydration planning and segment building
pack_objects,hydr_cover,hydr_pendingSQLite tableslastPackOids,lastPackKey,packListDO KV keyshasLooseBatch(),getObjectStream(),getObjectSize(),getPackLatest(),getPackOids(),getPackOidsBatch()RPCsRepoStorageModetype and all storage-mode switching- Rollback backfill machinery
isomorphic-gitdependency- Configuration:
REPO_UNPACK_*,REPO_KEEP_PACKS,REPO_PACKLIST_MAX,REPO_DO_MAINT_MINUTES
Module Deliverables
The implementation must keep repoDO.ts thin by pushing logic into helper modules.
Required new modules:
src/do/repo/catalog.tssrc/do/repo/receiveLease.tssrc/do/repo/compactionLease.tssrc/do/repo/legacyCompat.tssrc/git/receive/pktSectionStream.tssrc/git/receive/streamReceivePack.tssrc/git/pack/indexer/scan.tssrc/git/pack/indexer/inflateCursor.tssrc/git/pack/indexer/resolve.tssrc/git/pack/indexer/writeIdx.tssrc/git/object-store/catalog.tssrc/git/object-store/idxView.tssrc/git/object-store/lookup.tssrc/git/object-store/readObject.tssrc/git/object-store/readRefsBatch.tssrc/git/object-store/delta.tssrc/git/compaction/plan.tssrc/git/compaction/run.tssrc/git/compaction/legacyLoosePack.ts
Primary files to update:
src/routes/git.tssrc/routes/admin.tssrc/routes/ui/adminPage.tssrc/routes/ui/overview.tssrc/routes/ui/tree.tssrc/routes/ui/commits.tssrc/routes/ui/raw.tssrc/routes/ui/helpers.tssrc/index.tssrc/common/progress.tssrc/do/repo/repoDO.tssrc/do/repo/debug.tssrc/do/repo/db/schema.tssrc/do/repo/db/dal.tssrc/client/components/ProgressBanner.tsxsrc/client/pages/AdminPage.tsxsrc/client/pages/CommitsPage.tsxsrc/client/pages/OverviewPage.tsxsrc/client/pages/TreePage.tsxsrc/client/islands/repo-admin/types.tssrc/client/islands/repo-admin/useRepoAdminActions.tssrc/client/islands/repo-admin/RepoOverviewCard.tsxsrc/client/islands/repo-admin/HydrationCard.tsxsrc/client/islands/repo-admin/PackFilesCard.tsxsrc/client/islands/repo-admin/index.tsxsrc/git/operations/closure.tssrc/git/operations/fetch/neededFast.tssrc/git/object-store/catalog.tssrc/git/operations/read/diff.tssrc/git/operations/read/objects.tssrc/git/operations/read/tree.tssrc/git/operations/read/commits.tssrc/git/pack/assemblerStream.tswrangler.jsonc
Phased Implementation Plan (completed)
Archival: All phases are complete and the closure release has removed all legacy code paths. This section is preserved as historical rollout context. See
MIGRATION-STREAMING-PUSH.mdfor the current operator guide.
Phase 1: Pack Catalog and Worker Object Store Shadow Mode (completed)
Added pack_catalog schema and DAL helpers, DO helpers for catalog reads/leases/legacy seeding, worker-local IdxView and packed-object resolver, read-path shadow validators, and queue binding. Automatic pack_catalog backfill for packed repos.
Phase 2: Read Path Cutover (completed)
Switched fetch planning, object reads, and all UI routes to worker-local pack-first object store. Removed correctness dependence on DO getObject*, hasLooseBatch, getPackOids*, and getObjectRefsBatch. Replaced unpack-progress UI with repo-activity UI. Added admin-only storage-mode control for canary validation.
Phase 3: Streaming Receive Cutover (completed)
Added receive lease RPCs, staged R2 write, pack indexer, thin-pack validation, connectivity checks, and DO finalization. Kept packList and lastPackKey mirrored for rollback. Added emergency legacy backfill tooling for canary rollback.
Phase 4: Queue-Driven Compaction (completed)
Implemented beginCompaction / commitCompaction / abortCompaction and queue consumer. Tiered compaction of active catalog. Repurposed /admin/hydrate as compaction alias and added /admin/compact.
Phase 5: Legacy Path Retirement (completed)
Removed legacy receive/unpack/hydration methods from active code paths. Moved all repos to streaming. Removed legacy RPCs and tests. The closure release then removed all remaining legacy code, schema, and configuration.
Automated Testing Plan
Automated tests must not assume uncommitted-fixture/ is committed.
New or rewritten worker tests
- Receive path streams without
request.arrayBuffer(). - Single active receive lease returns
503to concurrent pushes. - Thin-pack base resolution uses existing packed objects, not loose objects.
- Fetch works after deleting all DO
obj:*. - Fetch works after deleting all R2 loose mirrors.
hasObjectsBatch()andreadObjectRefsBatch()are pack-first.- Compaction replaces source packs and keeps fetch correct.
DELETE /admin/pack/:packKeyrejects deletion of active packs.- Debug endpoints continue to work after all loose data is removed.
- Test repo seed helpers can build pack-only repos without writing loose objects.
- Pack-first test seeding replaces
seedMinimalRepo()loose assumptions in new and migrated tests.
Legacy tests (removed in closure)
The following test categories were deleted in the closure release:
- hydration clear/delete tests
- unpack progress tests
- one-deep unpack queue tests
pack_objectsexact membership tests- storage-mode transition tests
- rollback compatibility tests
Performance and memory tests
Add a local-only benchmark script under scripts/ that:
- uses
uncommitted-fixture/when present; - measures peak memory, index time, and receive latency;
- is not required for CI.
CI should instead generate smaller deterministic packs with:
- long OFS-delta chains;
- thin REF_DELTA bases;
- multiple receive packs triggering compaction.
Test seeding must also change:
- stop relying on
seedMinimalRepo()writing loose objects for packed repos; - add a packed-only seed helper and use it in all fetch/read tests that are meant to validate the new correctness path.
Acceptance Criteria
The rollout is successful only if all of the following are true:
- A repo push with about 8,000 commits, about 70,000 loose objects’ worth of content, a 40 MB pack, and a 3 MB idx completes without buffering the whole pack in memory.
- After such a push, deleting all DO
obj:*keys does not break fetch, tree/blob/raw reads, diff, or commit browsing. - After such a push, deleting all R2 loose mirrors does not break fetch, tree/blob/raw reads, diff, or commit browsing.
- Fetch planning no longer depends on
pack_objects,lastPackOids,unpackWork,unpackNext, or hydration state. - Active pack count remains bounded by the compaction policy after repeated pushes.
- The DO is never required to materialize object data for correctness.
- The only metadata authority remains the DO.
- No route order or auth behavior regresses.
- Existing packed repos migrate without requiring loose objects.
- Legacy loose-only repos have an explicit migration path before cutover.
Final Notes For Implementers
- Do not bloat
src/do/repo/repoDO.ts. Every new behavior belongs in a helper module and is only wired through thin RPC methods. - Do not add raw Drizzle queries outside the DAL.
- Do not reintroduce exact per-pack SQLite membership as a correctness dependency.
- Do not make loose objects required again, even as a “temporary shortcut”.
- Do not store repo metadata in R2.
- Prefer simple fixed constants over new env vars unless a hard operational need appears during implementation.
Closure Release Implementation Checklist (completed)
All items below were completed in the closure release. See docs/streaming-push-closure-plan.md for the full implementation plan.
Remove emergency backfill tool-- doneRemove-- done/admin/hydratealiases andstartHydration/clearHydrationRPCsRemove-- donemirrorLegacyPackKeys()andlastPackKey/lastPackOids/packListfromRepoStateSchemaRemove-- doneREPO_UNPACK_*fromwrangler.jsonc,repoConfig.ts,vitest.bindings.tsDelete legacy receive, unpack, hydration code-- doneRemove-- doneunpackWork,unpackNext,hydrationWork,hydrationQueuefromRepoStateSchemaDrop-- donepack_objects,hydr_cover,hydr_pendingtablesRemove-- doneisomorphic-gitfrompackage.jsonRemove-- doneRepoStorageModetype and all storage-mode RPCs/admin endpointsDelete remaining legacy tests-- doneFinal docs cleanup-- done