Blob: docs/better-fetch.md
Modernize Fetch/Upload-Pack: Ordered Pack Snapshot + Shared Rewrite Core
Context
Phase 3 of the streaming-push work already introduced the right low-level building blocks:
IdxViewwith typed arrays and byte-bounded caching- pack-first object lookup and object reads in Worker code
- the newer streaming receive pipeline and pack indexer
- iterative, bounded-memory processing instead of large ad hoc Maps
The fetch/upload-pack path still predates that work. It still has its own idx parsing stack, its own pack metadata layer, multiple fetch-only serving modes, and compatibility fallback branches that should no longer be part of fetch correctness.
This proposal replaces the current fetch assembler stack with one ordered pack-snapshot rewrite pipeline that:
- serves fetch immediately
- stays correct when loose compatibility data is deleted
- reuses the pack-first object store
- is shaped so Phase 4 compaction can reuse the same rewrite core
No new schema, sidecar metadata, or env vars are introduced.
Pack Ordering Contract
The rewrite engine receives a caller-ordered snapshot and must honor that order for duplicate-object selection and stable tie-breaking.
- Fetch passes active packs in the exact order returned by the active pack catalog snapshot.
- Today that order is
seqHi DESC, tier DESCfor active rows. - Phase 4 compaction will pass
[source packs..., remaining active packs...].
This contract matters because duplicate selection must stay deterministic and because compaction needs explicit control over which pack wins when the same object appears in multiple sources.
Reader Reuse Decision
Do not invent a third unrelated pack-read strategy.
The new rewrite engine should follow the same buffered range-read policy already used by the newer pack-indexer resolve reader:
- preload sequential chunks when locality exists
- provide direct
readRange()for exact spans - provide
readWindow()for progressive reads without double buffering
In this fetch refactor, the rewrite path should start with that policy instead
of the current fetch-only groupCache design.
If measurements later show that fetch needs a stronger byte-budgeted multi-window cache, add it as a narrow optimization after correctness is landed. Do not make that extra cache part of the first correctness rewrite.
Goals
- Do more with less code.
- Delete duplicated idx parsing and fetch-only pack metadata structures.
- Keep protocol v2 behavior, route behavior, and response semantics stable.
- Preserve request limiter and subrequest accounting.
- Keep Worker to DO to one metadata hop for the active catalog snapshot, then Worker to R2 only for fetch serving.
- Make fetch correctness depend only on the active pack catalog plus R2 packs.
- Produce one serving core that Phase 4 compaction can reuse.
Non-Goals
- Do not change Git protocol semantics.
- Do not restore loose-only fetch.
- Do not redesign UI or unrelated read paths in this pass.
- Do not remove legacy receive or hydration systems beyond the narrow migrations
required by idx-cache deletion and
IdxViewadoption. - Do not rename
readLooseObjectRaw()in this pass. - Do not add new database tables, catalog fields, or storage-mode flags.
Current Problems
The fetch path still carries multiple plan and serving branches:
InitCloneUnionIncrementalSingleIncrementalMulti- dead buffered-mode leftovers
The current assembler duplicates work that already has a better replacement:
src/git/pack/idxCache.tssrc/git/pack/packMeta.ts:parseIdxV2()- per-pack
Map<string, number>andMap<number, number>state insrc/git/pack/assemblerStream.ts
Fetch closure planning still has compatibility fallback reads:
- mainline enrichment in
src/git/operations/fetch/neededFast.ts - missing-ref fallback in
src/git/operations/fetch/neededFast.ts findCommonHaves()fallback insrc/git/operations/closure.ts
- mainline enrichment in
Initial clone and closure-timeout fallback paths can inflate fetch scope from the actual closure result to a full pack union.
src/git/object-store/store.ts:readObjectRefsBatch()still walks objects serially even though the new object-store path is otherwise pack-first.Hydration still depends on the old parsed-idx shape in:
src/do/repo/hydration/status.tssrc/do/repo/hydration/stages/scanDeltas.ts
The current assembler shape is fetch-specific and not a good Phase 4 compaction surface.
Files To Delete
| File | Why |
|---|---|
src/git/pack/idxCache.ts |
Duplicate idx cache; loadIdxView() already provides the needed cache and request-local coalescing |
src/git/pack/assemblerStream.ts |
Replaced by rewrite.ts |
src/git/operations/heavyMode.ts |
Only exists to shape compatibility loose fallback behavior during closure |
Files To Create
| File | Purpose |
|---|---|
src/git/pack/rewrite.ts |
Shared pack rewrite engine for fetch now and compaction later |
Core Types
type OrderedPackSnapshot = {
packs: {
packKey: string;
packBytes: number;
idx: IdxView;
}[];
};
type UploadPackPlan =
| {
type: "Serve";
repoId: string;
snapshot: OrderedPackSnapshot;
neededOids: string[];
ackOids: string[];
signal?: AbortSignal;
cacheCtx?: CacheContext;
}
| {
type: "RepositoryNotReady";
};This replaces InitCloneUnion, IncrementalSingle, and IncrementalMulti.
There should be no public serveMode field. Fast-path passthrough decisions
belong inside the rewrite engine after it has already selected the actual object
set.
High-Level Flow After Refactor
handleFetchV2Streaming()parses the request and keeps the current sideband response shape.The planner loads the active catalog once and eagerly resolves an ordered snapshot with
packBytesandIdxView.The planner computes:
ackOidsfrom pack-first membership checksneededOidsfrom pack-first closure only- initial-clone unions directly from eager
IdxViews
resolvePackStream()calls one rewrite entrypoint with the ordered snapshot andneededOids.The rewrite engine emits a valid PACK stream:
- select needed objects from ordered source packs
- pull in delta bases
- topologically order base before dependent
- rewrite OFS distances until stable
- stream payload bytes from R2 with bounded memory
Detailed File Plan
1. src/git/operations/fetch/types.ts
Replace the current plan union with OrderedPackSnapshot plus UploadPackPlan.
Required changes:
- remove
InitCloneUnion,IncrementalSingle, andIncrementalMulti - add
OrderedPackSnapshot - add
UploadPackPlan - keep
ackOids,signal, andcacheCtxon the serve plan
Why:
- the old branches are implementation artifacts, not meaningful protocol states
- a single serve plan makes fetch and compaction share the same serving contract
2. src/git/operations/fetch/plan.ts
Rewrite the planner around OrderedPackSnapshot.
Required changes:
Load active catalog rows directly with
loadActivePackCatalog().If the active catalog is empty, return
RepositoryNotReady.For each active row:
- call
loadIdxView(env, packKey, cacheCtx, packBytes) - build one snapshot entry with
packKey,packBytes, andidx
- call
Keep active catalog order as the authoritative pack preference order.
Reuse request-local memo state so later object-store calls see the same
packCatalogandidxViews.Do not route the serving plan through
getPackCandidates().
Implementation footgun:
The planner is the one place that has both packKey and packBytes. If it
falls back to getPackCandidates(), later loadIdxView() calls lose the size
hint and pay avoidable head() reads.
Initial clone path:
- if
haves.length === 0, enumerate the union directly from eagerIdxViews - iterate
idx.countand materialize OIDs withgetOidHexAt() - deduplicate with
Set<string> - delete
buildUnionNeededForKeys() - delete
countMissingRootTreesFromWants()
Incremental path:
- call
computeNeededFast()using only the pack-first object store - remove
beginClosurePhase()andendClosurePhase() - if closure times out, return the partial
neededOidsresult instead of upgrading to a full union
Negotiation:
- compute
ackOidswithfindCommonHaves()only when needed - keep the current
done ? [] : ackOidsbehavior
Logging:
- keep the current summary style
- add one snapshot summary log with pack count, idx loads, cheap hit/miss counters, and total indexed object count when cheap to compute
3. src/git/operations/closure.ts
Trim this module to the parts fetch still needs.
findCommonHaves():
- keep the 128-have cap
- keep the same return shape
- remove the fallback loop that calls
readLooseObjectRaw() - rely only on
hasObjectsBatch()
Delete:
buildUnionNeededForKeys()countMissingRootTreesFromWants()
Potential follow-up deletion:
iterPackOids()if nothing else still uses it
Why:
- these helpers only exist to support old fetch planning branches
- fetch negotiation should no longer cross into compatibility object reads
4. src/git/operations/fetch/neededFast.ts
Convert closure planning to fully pack-first behavior.
Required changes:
- Remove the compatibility fallback block that calls
readLooseObjectRaw()for missing refs. - Remove
heavyModeintegration andloader-cappedhandling. - Keep timeout flagging so callers can still detect partial closure.
- Preserve the stop-set and mainline optimization, but make mainline enrichment pack-first.
Mainline enrichment:
- replace
readLooseObjectRaw()with pack-first reads - reuse the object store, not a second ad hoc reader
- preserve the current guard and budget shape unless it is intentionally
remeasured:
- only run when
ackOids.length > 0 && ackOids.length < 10 - max 20 mainline steps
- stop after about 2 seconds
- only run when
Batching and memoization:
- keep queue batching at 128 logical OIDs
- keep request-local refs memo behavior
- continue omitting missing objects from the refs map rather than throwing
Failure behavior:
- missing objects remain omitted
- closure does not fall back to compatibility loose data
- timeout returns the partial closure result already accumulated
Why:
- this removes the last fetch correctness dependence on compatibility reads
5. src/git/object-store/store.ts
Make readObjectRefsBatch() bounded-parallel.
Required changes:
- process OIDs in bounded batches using
MAX_SIMULTANEOUS_CONNECTIONS - within each batch, use
Promise.all()overreadObject() - preserve current semantics:
- missing objects remain omitted
- commit/tree/tag/leaf handling stays the same
- request-local memoization still comes from
readObject()
Why:
- the current serial walk is a bottleneck on large closure traversals
6. src/git/pack/rewrite.ts
Create the shared rewrite engine that replaces assemblerStream.ts.
Public entrypoint:
export async function rewritePack(
env: Env,
snapshot: OrderedPackSnapshot,
neededOids: string[],
options?: {
signal?: AbortSignal;
limiter?: { run<T>(label: string, fn: () => Promise<T>): Promise<T> };
countSubrequest?: (n?: number) => void;
onProgress?: (msg: string) => void;
}
): Promise<ReadableStream<Uint8Array> | undefined>;Internal design:
- use the ordered snapshot directly
- do not reload pack metadata through a second fetch-only stack
- keep selection and layout state in typed arrays
- allow small targeted lookup state where it improves readability, but do not
rebuild full per-pack OID and offset Maps that
IdxViewalready replaced
Suggested per-selection state:
const packSlot = new Uint8Array(capacity);
const entryIndex = new Uint32Array(capacity);
const offset = new Float64Array(capacity);
const nextOffset = new Float64Array(capacity);
const typeCode = new Uint8Array(capacity);
const origHeaderLen = new Uint16Array(capacity);
const baseSlot = new Int32Array(capacity);
const sizeVarBuf = new Uint8Array(capacity * 5);
const sizeVarLen = new Uint8Array(capacity);
const outHeaderLen = new Uint16Array(capacity);
const outOffset = new Float64Array(capacity);Rewrite Phases
Select objects.
- for each
neededOid, search packs in snapshot order withfindOidIndex() - first pack wins on duplicates
- deduplicate by
(packSlot, entryIndex), not by OID string - include all required delta bases
- for
OFS_DELTA, resolve base by offset within the same pack - for
REF_DELTA, resolve base OID by searching the ordered snapshot
- for each
Detect passthrough.
After selection, if:
- the snapshot has exactly one pack
- every object in that pack is selected
- no header rewrite is needed
then stream the existing
.packbody from R2, strip the old 20-byte trailer, hash the bytes as they stream, and emit a fresh trailer.Keep this decision inside the engine, not in the planner.
Read headers in batches.
- sort selected entries by
(packSlot, offset) - coalesce header reads instead of fetching headers one at a time
- whole-pack preload remains valid for small packs
- reuse
readPackHeaderExFromBuf()when whole-pack bytes are already in memory
- sort selected entries by
Topologically order output.
- base before dependent
- stable tie-break by snapshot order first, then source offset
- on incomplete ordering or cycle, log and fail the request
Converge output header lengths.
- preserve compressed payload bytes unchanged
- recompute OFS distances with
encodeOfsDeltaDistance() - iterate until header lengths stabilize or a fixed sanity cap is reached
- if convergence fails, log and fail the request
Stream the output.
- emit PACK header
- emit rewritten entry headers and original compressed payloads
- emit SHA-1 trailer
- keep sideband outside the rewrite engine
Read Policy Inside rewrite.ts
Do not reintroduce the old fetch-only groupCache as the default design.
Instead:
- keep a per-pack buffered reader with the same chunked preload and
readWindowbehavior already proven in the newer pack-indexer resolve path - use exact
readRange()for large spans - keep whole-pack preload for small packs
If a second-stage optimization is needed after correctness lands, add a byte-budgeted per-pack window cache and measure it explicitly. That should be an optimization patch, not part of the first correctness rewrite.
Limiter and subrequest accounting:
- every R2 read path inside the engine must respect the request limiter
- every R2 read path must account through
countSubrequest() - preserve the existing one-shot soft-budget warning style
Why:
- this gives fetch and compaction one serving core
- it deletes duplicated idx parsing and most fetch-only metadata code
- it stays aligned with the newer streaming push patterns
7. src/git/operations/fetch/execute.ts
Collapse this module to a thin wrapper around the rewrite engine.
Required changes:
- delete single-pack versus multi-pack dispatch
- delete fallback-from-single-to-multi retry
- call
rewritePack(env, plan.snapshot, plan.neededOids, options)
Why:
- the planner already has the ordered snapshot
- the rewrite engine should be the only serving path
8. src/git/operations/uploadStream/index.ts
Keep this as the protocol entrypoint, but simplify it.
Required changes:
- remove the early
getPackCandidates()preflight - remove the re-export of
computeNeededFast - keep:
- request parsing
- negotiation-only response path
- sideband muxing
- fatal response behavior
499handling on abort
Behavior to preserve:
- no buffered mode
- no
X-Git-Streamingswitching - same
repositoryNotReadyResponse()behavior when no active packs exist
9. src/git/operations/fetch/protocol.ts
Delete the dead buffered helper:
respondWithPacketizedPack()
Keep:
buildAckSection()buildAckOnlyResponse()
Update tests that still import the deleted helper.
10. src/git/pack/packMeta.ts
Trim this module to low-level pack helpers still used elsewhere.
Keep:
readPackHeaderEx()readPackHeaderExFromBuf()readPackRange()encodeOfsDeltaDistance()mapWithConcurrency()if it still has callers
Delete:
IdxParsedPackMetaloadPackMeta()parseIdxV2()readUint64BE()if it becomes unused
Why:
IdxViewis the shared idx representation now- these legacy shapes only exist to support the old fetch assembler and old hydration callers
11. Hydration callers
Migrate off loadIdxParsed() in:
src/do/repo/hydration/status.tssrc/do/repo/hydration/stages/scanDeltas.ts
Required changes:
- replace
loadIdxParsed()withloadIdxView() - update
buildPhysicalIndex()so it can work from a narrowIdxView-based input - use
getOidHexAt()only where hex materialization is actually needed - use
findOidIndex()forREF_DELTAlookup instead of rebuilding full OID arrays for search
Important:
This is not a purely mechanical swap. buildPhysicalIndex() currently expects
plain string-array OIDs and offset arrays, so the helper needs to be reshaped.
12. src/git/pack/index.ts
Update exports carefully.
Required changes:
- export
rewrite.tsinstead ofassemblerStream.ts - keep current exports that still have live callers:
unpack.tsloose-loader.tspackMeta.tsbuild.tsindexer/index.ts
Do not drop unpack.ts or loose-loader.ts from the barrel in this pass.
13. File deletions after cutover
Once the rewrite engine is wired in and validated, delete:
src/git/pack/assemblerStream.tssrc/git/pack/idxCache.tssrc/git/operations/heavyMode.ts
Deleted Or Simplified Paths
After the refactor, remove or simplify:
src/git/pack/assemblerStream.tssrc/git/pack/idxCache.tssrc/git/operations/heavyMode.tsrespondWithPacketizedPack()buildUnionNeededForKeys()countMissingRootTreesFromWants()- fetch-time compatibility fallback in
findCommonHaves() - fetch-time compatibility fallback in
computeNeededFast()
Potential follow-up deletion if it becomes unused:
iterPackOids()
Logging
Follow existing logging style:
- compact summary counters
- one-shot soft-budget warnings
- no logging tests
Useful new summary points:
snapshot load:
- pack count
- idx loads
- cheap cache hits versus misses
- total indexed objects when cheap
rewrite selection:
- requested OIDs
- selected entries
- added delta bases
streaming:
- whole-pack hits
- buffered-reader cache hits versus misses if cheap
- direct range fallbacks
- total time
Test Plan
Existing tests to keep green
- negotiation and ack behavior
- streaming pack response
- gzip request body handling
- fetch while receive or unpack work overlaps
- pack-first fetch after loose deletion
Existing tests to update
test/fetch-streaming.worker.test.ts- remove the buffered-mode test
- remove
X-Git-Streamingfrom remaining fetch tests
test/upload-pack-acks.test.ts- stop importing
respondWithPacketizedPack()
- stop importing
test/pack-indexer.resolve.ofs.worker.test.ts- switch from
parseIdxV2()toparseIdxView()plusgetOidHexAt()
- switch from
tests that import
computeNeededFastfromuploadStream/index.ts- import from
fetch/neededFast.tsdirectly if they still need the symbol
- import from
New tests to add
Single-pack fetch with all loose objects deleted.
Multi-pack incremental fetch where required bases span multiple packs.
Duplicate-object selection honoring active catalog order.
Cross-pack
REF_DELTArewrite correctness.OFS_DELTArewrite correctness with header-length convergence.Single-pack passthrough fast path.
Mid-stream abort preserving current
499and sideband fatal behavior.Closure timeout returning the partial closure result instead of silently upgrading to full union.
Rewrite contract for future compaction:
- ordered source packs first
- remaining active packs second
- only required external bases included
Differential correctness gate while both engines exist:
- run the new
rewrite.tsand oldassemblerStream.tson the same logical inputs - compare pack validity and object coverage
- use this as a temporary landing gate before deleting the old assembler
Validation Commands
Minimum validation for this refactor:
npm run typecheck
npx vitest run --config vitest.config.ts test/fetch-streaming.worker.test.ts
npx vitest run --config vitest.config.ts test/pack-first-fetch-and-ui.worker.test.ts
npx vitest run --config vitest.config.ts test/fetch-during-unpack.worker.test.ts
npx vitest run --config vitest.config.ts test/pack-first-read-path.closure.worker.test.ts
npx vitest run --config vitest.config.ts test/upload-pack-content-encoding.worker.test.tsAdditional targeted checks:
npx vitest run --config vitest.config.ts test/receive-push.worker.test.ts
npx vitest run --config vitest.config.ts test/streaming-receive.worker.test.ts
npx vitest run --config vitest.config.ts test/pack-indexer.resolve.ofs.worker.test.ts
npx ava test/object-parse.test.ts
npx ava test/ofs-delta-encode.test.ts
npx ava test/ofs-delta-known-encodings.test.tsFull validation before landing:
npm run typecheck
npm run test
npm run test:workers
npm run format:checkRecommended Implementation Order
- Introduce
OrderedPackSnapshotand the simplified serve-plan types. - Rewrite the planner to load the active catalog and eager
IdxViews directly, but keep the current assembler temporarily so the surface area shrinks first. - Remove closure compatibility fallbacks and make
readObjectRefsBatch()bounded-parallel. - Implement
src/git/pack/rewrite.ts. - Add a temporary differential test that compares the new rewrite path against
assemblerStream.tson the same inputs while both still exist. - Wire
resolvePackStream()to the new rewrite engine. - Delete
assemblerStream.ts,idxCache.ts,heavyMode.ts, and the dead buffered helper. - Trim
packMeta.ts. - Migrate hydration callers to
IdxView. - Update tests and run full validation.
Acceptance Criteria
- Fetch correctness depends only on the active pack catalog plus R2 packs.
- No fetch correctness path depends on compatibility loose objects.
- The planner loads the active catalog once and each idx once per request.
- The planner does not route serving-path planning through
getPackCandidates(), becausepackBytesmust be preserved for hintedloadIdxView()calls. - Single-pack and multi-pack fetch use the same rewrite engine.
- Duplicate-object choice follows active catalog order.
- Mainline enrichment in
computeNeededFast()keeps the current guard and budget shape unless it is intentionally remeasured and changed. - The rewrite engine is reusable for Phase 4 compaction by passing a different ordered snapshot.
- Dead fetch-specific idx parsing and buffered-mode code is removed.
- The new rewrite path uses the same buffered pack-read policy as the newer pack-indexer resolve path, unless later measurement proves a stronger cache is necessary.