Skip to content
File

Blob: docs/REPO_ACCESS_PERSISTENCE_PLAN.md

Markdown504 lines

Repo Access And Persistence Plan

This document proposes the least painful integration path for:

  • private GitHub repository access without exposing long-lived credentials inside the sandbox
  • R2-backed backup and restore for /workspace
  • a multi-tenant metadata layer that fits the existing user/workspace -> sandbox ownership model

It is based on the current ccccocc codebase and current official Cloudflare/GitHub documentation as of April 14, 2026.

Current repo constraints

The live tree already has the right base model:

  • Cloudflare Access proves identity, and the Worker derives sandbox ownership from userId + workspace in src/worker/auth.ts and src/worker/index.ts.
  • The app is terminal-first. The persisted browser model stores logical tabs and backend session IDs, not repo metadata, in src/client/workspace/store.ts.
  • Backup and restore routes already exist in src/worker/index.ts, but wrangler.jsonc does not yet bind an R2 backup bucket or related secrets.
  • Session env vars are intentionally sanitized server-side, including token-like names, so repo credentials should not be passed through createSession().

These constraints are good. They push the design toward Worker-held credentials and workspace-scoped metadata instead of shell env secrets.

Recommendation

Use this default architecture:

  1. Keep Cloudflare Access as the front-door identity layer only.
  2. Store per-workspace repo credentials and backup handles outside the sandbox.
  3. Use Worker-side outbound Git auth injection for github.com so terminal git clone, git fetch, and git pull can work without exposing the PAT to the sandbox.
  4. Use R2-backed createBackup() and restoreBackup() for /workspace.
  5. Use D1 as the authoritative multi-tenant metadata store.
  6. Add a small workspace-scoped coordination layer only if restore/backup concurrency becomes a real problem.

Inferred recommendation from the sources:

  • D1 should be the default metadata system.
  • A separate coordination Durable Object is optional, not required, and should be added only for serialized restore/backup workflows or other workspace-level locks.
  • KV should not be the authoritative metadata store for multi-tenant backup or credential state because KV is eventually consistent.

Why this path

Cloudflare now supports Worker-side outbound interception for sandbox HTTP and HTTPS traffic, including secure credential injection and Git-specific auth helpers. This matches ccccocc better than injecting a PAT into the shell.

Cloudflare's backup API already stores filesystem snapshots in R2 and explicitly recommends persisting backup handles externally in KV, D1, or Durable Object storage. That means the missing piece in this repo is not backup capability, but persistent metadata and UI/API plumbing.

Private repo access

Preferred path: outbound Git auth injection

Use an outbound handler on the sandbox class for github.com and let the Worker attach auth at request time.

What changes:

  • Replace export { Sandbox } from "@cloudflare/sandbox" with a local subclass in src/worker/index.ts.
  • Export ContainerProxy from the Worker entrypoint. Cloudflare documents this as required for outbound interception.
  • Define a named outbound handler such as authenticatedGithub.
  • Use sandbox.setOutboundByHost("github.com", "authenticatedGithub") for sandboxes that are allowed to access a private repo.
  • Keep the PAT in Worker-controlled storage, not in sandbox env vars and not embedded in the clone URL.

This gives the best operator experience because:

  • git clone, git fetch, and git pull work from inside the terminal.
  • token rotation happens in the Worker metadata layer
  • the sandbox never receives the raw PAT

Cloudflare's outbound guide shows the relevant pattern directly with authenticateGitHttpsRequest(request, githubToken, ctx.containerId).

Alternate path: one-shot checkout endpoint

If you want the smallest first step, add a POST /api/repos/checkout route that:

  • authenticates the caller with Access
  • loads the workspace PAT from metadata
  • checks out the repo into /workspace/...
  • returns { targetDir, repoUrl, branch }

This is enough for "connect repo" onboarding, but it does not solve normal terminal Git usage by itself. If the user later runs git fetch in the terminal, that terminal still needs Git auth. Because of that, one-shot checkout is a valid Phase 1, not the end state.

PAT handling

For the PAT-based path:

  • accept only fine-grained PATs for the MVP
  • store them encrypted at rest in the Worker metadata layer
  • never echo them back to the client after save
  • never place them in sessionStorage, query params, or session env

Recommended secret handling:

  • Worker secret: WORKSPACE_SECRET_KEY
  • encryption: AES-GCM using Web Crypto
  • persisted fields: ciphertext, iv, key_version, updated_at

If the product later becomes org-facing or needs delegated installs, replace PAT storage with a GitHub App flow. The rest of this architecture still holds.

Backup and restore

Required runtime changes

To make the existing backup routes actually work in production, add:

  • an R2 bucket binding BACKUP_BUCKET
  • a D1 binding such as METADATA_DB
  • Worker vars:
    • BACKUP_BUCKET_NAME
    • CLOUDFLARE_ACCOUNT_ID
  • Worker secrets:
    • R2_ACCESS_KEY_ID
    • R2_SECRET_ACCESS_KEY

wrangler.jsonc needs:

  • r2_buckets binding for BACKUP_BUCKET
  • d1_databases binding for METADATA_DB
  • vars.BACKUP_BUCKET_NAME
  • vars.CLOUDFLARE_ACCOUNT_ID

After changing config:

  • run npm run types:generate
  • run npm run typecheck

Backup handle persistence

Cloudflare's backup guide is explicit that DirectoryBackup handles are serializable and should be stored externally for later restore.

That means the app needs metadata for each backup:

  • owning userId
  • workspace
  • backup handle JSON
  • backup id
  • name
  • dir
  • ttl
  • whether useGitignore was used
  • createdAt
  • expiresAt
  • restoredAt
  • optional deletedAt

The current POST /api/workspace/backup and POST /api/workspace/restore routes are good primitives, but they are not enough for UI or multi-tenant restore by themselves because they assume the caller already knows the backup handle.

Restore semantics

Cloudflare's docs recommend stopping writes before restoring.

That matters in ccccocc because one workspace can have multiple logical terminal tabs and multiple backend sessions writing to the same /workspace.

Recommended restore behavior:

  • treat restore as a workspace-wide operation, not a per-tab action
  • block or warn when sessions are currently attached
  • after a restore, force-reset all workspace sessions or detach/reconnect them to avoid stale shell state

This is the main place where explicit coordination may be needed.

Metadata layer options

KV

KV is the simplest store for prototypes and single-record lookups, but Cloudflare documents KV as eventually consistent.

Use KV only for:

  • non-authoritative caches
  • transient lookup helpers
  • fallback bootstrap state

Do not use KV as the source of truth for:

  • current workspace PAT
  • backup list UI
  • restore eligibility
  • last-restored state
  • concurrency-sensitive operations

D1

D1 is the recommended authoritative metadata store for this project.

Why D1 fits:

  • SQL queries are a better fit for listing backups, filtering by workspace, and auditing state
  • strong enough semantics for app metadata without adding actor-style complexity everywhere
  • easy to bind into the existing Worker
  • Cloudflare explicitly positions D1 for per-user, per-tenant, or per-entity databases

Recommended D1 usage in ccccocc:

  • one small metadata database for the app at first
  • tables keyed by user_id and workspace
  • indexes for user_id, workspace, backup_id, and active credential records

Suggested initial schema:

CREATE TABLE workspace_integrations (
  user_id TEXT NOT NULL,
  workspace TEXT NOT NULL,
  provider TEXT NOT NULL,
  auth_mode TEXT NOT NULL,
  token_ciphertext BLOB,
  token_iv BLOB,
  key_version TEXT NOT NULL,
  created_at TEXT NOT NULL,
  updated_at TEXT NOT NULL,
  PRIMARY KEY (user_id, workspace, provider)
);

CREATE TABLE workspace_repos (
  user_id TEXT NOT NULL,
  workspace TEXT NOT NULL,
  host TEXT NOT NULL,
  owner TEXT NOT NULL,
  repo TEXT NOT NULL,
  branch TEXT,
  target_dir TEXT NOT NULL,
  created_at TEXT NOT NULL,
  updated_at TEXT NOT NULL,
  PRIMARY KEY (user_id, workspace, host, owner, repo)
);

CREATE TABLE workspace_backups (
  user_id TEXT NOT NULL,
  workspace TEXT NOT NULL,
  backup_id TEXT NOT NULL,
  name TEXT,
  dir TEXT NOT NULL,
  ttl_seconds INTEGER NOT NULL,
  use_gitignore INTEGER NOT NULL DEFAULT 0,
  handle_json TEXT NOT NULL,
  created_at TEXT NOT NULL,
  expires_at TEXT NOT NULL,
  restored_at TEXT,
  deleted_at TEXT,
  PRIMARY KEY (user_id, workspace, backup_id)
);

CREATE INDEX idx_workspace_backups_lookup
  ON workspace_backups (user_id, workspace, created_at DESC);

Durable Object

Do not add a new metadata Durable Object by default.

Durable Objects are the right tool when you need:

  • single-threaded coordination
  • strong per-workspace serialization
  • alarms or background wakeups
  • in-memory state plus durable state in one place

That makes them a good fit for a narrow coordinator such as WorkspaceCoordinatorDO, keyed by userId:workspace, if and only if you need one of these:

  • prevent concurrent restore and backup operations
  • gate restore while sessions are attached
  • maintain per-workspace ephemeral state such as "restore in progress"
  • coalesce repeated backup requests

Recommended rule:

  • D1 stores authoritative metadata
  • optional DO enforces workspace-local locks

This keeps the repo aligned with ccccocc's existing "avoid extra control plane infrastructure" rule while still giving a clean escape hatch for coordination.

Multi-tenant model

The existing ownership model should remain unchanged:

  • Access identity proves userId
  • client sends workspace
  • Worker derives sandbox ownership as ${userId}-${workspace}

All metadata should use the same compound key:

  • user_id
  • workspace

Do not key metadata by raw sandbox ID alone. The sandbox ID is a derived runtime identifier. The authoritative app concept is still (user, workspace).

Recommended tenancy rules:

  • one PAT or GitHub integration record per user_id + workspace + provider
  • one repo catalog per user_id + workspace
  • many backup records per user_id + workspace
  • all read and write operations re-check Access identity before touching metadata

Frontend changes

The current frontend only models sessions and tabs. Add a separate workspace settings model for repo and persistence state.

Recommended UI additions:

  • GitHub access section
    • Connect PAT
    • Update PAT
    • Remove PAT
    • status only, never token reveal
  • Repository section
    • repo URL or owner/repo
    • branch
    • target directory under /workspace
    • shallow clone toggle
    • optional "open in new tab"
  • Snapshots section
    • create backup
    • list backups
    • restore
    • delete backup metadata entry
    • restore warning: workspace-wide action

The tab/session store should keep only non-secret workspace metadata such as:

  • active repo path
  • repo URL
  • default branch
  • last backup ID

It should not hold the PAT or backup handles.

Backend routes

Recommended new routes:

  • GET /api/integrations/github/status?workspace=...
  • POST /api/integrations/github/pat?workspace=...
  • DELETE /api/integrations/github/pat?workspace=...
  • POST /api/repos/checkout?workspace=...
  • GET /api/workspace/backups?workspace=...
  • DELETE /api/workspace/backups?workspace=...&backup=...

Recommended behavior changes to existing routes:

  • POST /api/workspace/backup
    • persist returned backup handle to D1
    • support name, ttl, and useGitignore
  • POST /api/workspace/restore
    • load handle from D1 by backup ID, not from an opaque client-provided object
    • record restored_at
    • optionally trigger workspace session resets after restore

Auth injection integration details

Initial implementation

Implement this in the Worker:

  1. Load the workspace PAT from D1 after Access auth succeeds.
  2. Decrypt the PAT inside the Worker.
  3. Enable sandbox outbound auth for github.com.
  4. Route Git HTTPS traffic through authenticatedGithub.

This keeps the credential outside the sandbox while still allowing terminal-native Git usage.

Credential lookup shape

Cloudflare's outbound docs show ctx.containerId in the outbound handler. They also show per-instance secret lookup by ctx.containerId.

Inference for ccccocc:

  • D1 should remain the source of truth keyed by user_id + workspace
  • if the outbound handler cannot directly derive workspace identity, maintain a small ephemeral mapping from active container identity to workspace credential reference
  • that ephemeral mapping can live in KV or a narrow coordination DO, but it should not replace D1 as the source of truth

This preserves multi-tenant correctness across container restarts while still fitting the outbound handler model.

Hardening

Future hardening options:

  • add allowedHosts and deny-by-default outbound rules if the product later wants stricter egress control
  • add separate host handlers for api.github.com only if the app begins to proxy GitHub API calls
  • add additional GitHub hosts only when required by concrete workflows such as LFS or raw-content fetches

Do not flip enableInternet = false as part of the first repo-access change unless the app has already inventoried every external hostname needed by Codex CLI, Claude Code, package managers, and other agent tooling.

R2-backed persistence vs mounted buckets

For this repo, backup/restore should come first.

Reasons:

  • it matches the current /workspace model directly
  • it is already partially implemented in the Worker
  • it avoids turning /workspace into an object-storage mount
  • Cloudflare's mount-bucket docs recommend avoiding /workspace as a mount path and note that mounted storage is slower than local filesystem

Mounted buckets remain useful later for:

  • shared datasets
  • explicit persistent directories like /data
  • cross-sandbox shared artifacts

They should not be the first persistence layer for terminal workspaces in this app.

Open questions for a spike

These points should be validated in a short implementation spike before the full rollout:

  • whether the outbound Git handler can derive enough stable identity directly, or whether ccccocc needs an explicit container-to-workspace credential lookup table
  • whether the current terminal/session reconnect flow should hard-reset sessions after restore or only on explicit user confirmation
  • whether private-repo setup should begin with one-shot checkout first or go directly to terminal-native Git auth injection
  • which exact GitHub hosts need handlers beyond github.com for the workflows you actually want to support

Rollout plan

Phase 1

  • Add R2 backup configuration to wrangler.jsonc
  • Add D1 metadata binding as METADATA_DB
  • Add D1 schema and migration
  • Persist backup handles on backup creation
  • Add backup listing and restore-by-ID

Phase 2

  • Add PAT save/status/remove routes
  • Encrypt PATs at rest
  • Add repo settings UI
  • Add one-shot checkout endpoint

Phase 3

  • Replace one-shot-only checkout with outbound Git auth injection
  • Export ContainerProxy
  • Subclass Sandbox
  • Add named outbound handler for github.com
  • Verify terminal-native git fetch/pull/clone

Phase 4

  • Add workspace-wide restore coordination
  • Force-reset or reconnect sessions after restore
  • Add cleanup jobs, retention UI, and lifecycle documentation

Phase 5

  • Revisit GitHub App auth if PATs become operationally painful or org-facing

Testing and validation

Required automated checks after implementation:

  • npm run typecheck
  • npm test
  • npm run build when changing bindings, routes, or Worker wiring

Required manual checks:

  • Access-authenticated user A cannot view or restore user B metadata
  • private repo checkout works without the PAT appearing in terminal env or shell history
  • terminal-native git fetch and git pull still work after initial setup
  • backup creation returns persisted metadata
  • restore blocks or warns while sessions are actively writing
  • restoring a backup resets or cleanly reconnects terminal sessions
  • container restart plus restore-from-handle works as expected

Decision summary

Use this stack:

  • Cloudflare Access for user identity
  • Worker-held encrypted PATs for GitHub auth
  • outbound Git auth injection for terminal-native private repo access
  • R2-backed createBackup() and restoreBackup() for /workspace
  • D1 as the authoritative metadata layer
  • optional workspace-scoped Durable Object only for restore/backup coordination

This is the smallest path that:

  • respects the current ccccocc ownership model
  • keeps credentials out of the sandbox
  • supports multi-tenant backup and restore
  • avoids adding a new control plane unless coordination pressure justifies it

Sources