Blob: OPERATOR.md
OPERATOR.md
This document is the operator runbook for bland v1 as it exists in the live repo today.
When this file conflicts with older docs, the source tree wins.
Purpose
- Deploy the production Worker safely.
- Verify the core user flows after deploy.
- Triage incidents against the current Cloudflare topology.
- Be explicit about current operational gaps instead of inventing recovery paths that do not exist in the repo.
Production Topology
Current production runtime from wrangler.jsonc:
| Component | Current value |
|---|---|
| Worker name | bland |
| Domain | https://bland.tools |
| D1 | bland-prod |
| R2 | bland-uploads, bland-sites |
| Queue | bland-tasks |
| Durable Objects | DocSync, WorkspaceIndexer |
| Workers AI | binding AI, default model @cf/google/gemma-4-26b-a4b-it |
| Rate limits | RL_AUTH, RL_API, RL_AI |
| Assets binding | ASSETS |
Request routing in the live Worker:
- Requests for the configured public Sites domain dispatch to the Sites router before
/api,/uploads, assets, and the SPA shell. GET /api/v1/*and/uploads/*go through the Hono app./parties/*goes through PartyServer /DocSync.- AI requests
POST /api/v1/workspaces/:wid/pages/:id/{rewrite,generate,summarize,ask}go through the Hono app, gated byRL_AIand member-only entitlements, and stream SSE back to the client. - Document navigations are Worker-first SPA shell responses from
ASSETS. - Direct asset requests are served from
ASSETS.
Data Authority
- D1 is authoritative for users, workspaces, memberships, page metadata, invites, shares, and upload metadata.
DocSyncDO-local SQLite is authoritative for persisted Yjs document snapshots.WorkspaceIndexerDO-local SQLite is authoritative for the derived FTS index.- R2 stores upload blobs and derived public Sites artifacts. It is not an authorization store.
- The search queue carries derived
index-pagework only. Lost queue messages can stale search, but they do not lose source-of-truth data. - D1 read-after-write across requests depends on bookmark propagation via
x-bland-d1-bookmark.
Required Runtime Config
Production requires these runtime values:
LOG_LEVELALLOWED_ORIGINSPUBLISHED_SITE_DOMAINJWT_SECRETTESSERA_OIDC_ISSUERTESSERA_OIDC_CLIENT_IDTESSERA_OIDC_CLIENT_SECRET
Notes:
JWT_SECRET,TESSERA_OIDC_CLIENT_ID, andTESSERA_OIDC_CLIENT_SECRETare secrets.SENTRY_DSNis optional and used for client-side Sentry only. Worker-side Sentry is not implemented.ALLOWED_ORIGINScontrols both HTTP CORS and WebSocket origin checks. Keep it exact.PUBLISHED_SITE_DOMAINcontrols public Sites host dispatch. Empty disables Sites.- Register one tessera redirect URI per allowed app origin, ending in
/api/v1/oidc/callback.
Optional AI backend selection (defaults to Workers AI in production):
BLAND_AI_MODE—workers-ai(default),openai-compat, ormock. Production usesworkers-ai;mockis for E2E only and must not be set in production.BLAND_AI_WORKERS_CHAT_MODEL,BLAND_AI_WORKERS_SUMMARIZE_MODEL— override the default Workers AI model (@cf/google/gemma-4-26b-a4b-itfor both). Leave unset to use defaults.BLAND_AI_OPENAI_ENDPOINT,BLAND_AI_OPENAI_API_KEY,BLAND_AI_OPENAI_CHAT_MODEL,BLAND_AI_OPENAI_SUMMARIZE_MODEL— only needed whenBLAND_AI_MODE=openai-compat. Treat the API key as a secret if used.
First-Time Setup
Use this section when standing up bland in a fresh Cloudflare account. Skip it for subsequent deploys — those use the Deploy Runbook below.
1. Cloudflare account prerequisites
- A Cloudflare account with Workers, Durable Objects, D1, R2, and Queues enabled.
npx wrangler logincompleted locally against that account.- A zone for the production domain (
bland.toolsin the live repo; adjustwrangler.jsoncroutesandALLOWED_ORIGINSif deploying under a different hostname).
2. Create the bindings
wrangler.jsonc references bindings by name but does not create them. Create each one before the first deploy.
D1:
npx wrangler d1 create bland-prodCopy the returned database_id into wrangler.jsonc under d1_databases[0].database_id.
R2:
npx wrangler r2 bucket create bland-uploads
npx wrangler r2 bucket create bland-sitesQueue:
npx wrangler queues create bland-tasksDurable Objects (DocSync, WorkspaceIndexer), rate-limit bindings (RL_AUTH, RL_API, RL_AI), and the Workers AI binding (AI) are created implicitly on the first wrangler deploy from the classes and config already in the repo. Workers AI is enabled per-account in the Cloudflare dashboard; no secret or extra binding setup is required to use the default model.
3. Register tessera OIDC client
Register bland as an OIDC relying party with tessera:
- Redirect URI:
https://bland.tools/api/v1/oidc/callback - Redirect URI for every other app origin in
ALLOWED_ORIGINS, such ashttps://docs.limic.dev/api/v1/oidc/callback - Scopes:
openid email profile - Flow: authorization code with PKCE
4. Set runtime config
Public vars live in wrangler.jsonc → vars. Secrets are set per-environment with wrangler secret put.
Secrets (prompts for the value):
npx wrangler secret put JWT_SECRET # 32+ random bytes, e.g. `openssl rand -base64 48`
npx wrangler secret put TESSERA_OIDC_CLIENT_ID
npx wrangler secret put TESSERA_OIDC_CLIENT_SECRET
npx wrangler secret put SENTRY_DSN # optional; leave unset to disable client SentryNotes:
LOG_LEVEL,ALLOWED_ORIGINS,PUBLISHED_SITE_DOMAIN, andTESSERA_OIDC_ISSUERare already inwrangler.jsoncvars. Change them there if needed and redeploy.JWT_SECRETmust never be logged or committed. Rotate by setting a new value and redeploying; all existing access tokens become invalid and clients re-authenticate via the refresh cookie flow.- If upgrading a password-era deployment, follow MIGRATION-OIDC.md before applying the password-column closure migration.
5. Attach the custom domain
wrangler.jsonc registers bland.tools as a custom domain route. Point the zone's DNS at Cloudflare and let the first deploy bind the route, or pre-create the custom domain in the Workers dashboard. Update the routes[0].pattern and ALLOWED_ORIGINS if deploying under a different hostname.
6. Apply remote D1 migrations
Run migrations against the production D1 before the first deploy so the initial deploy starts against a populated schema:
npm run db:migrate:remotenpm run deploy re-runs this every time, so later deploys don't need a separate migration step.
7. First deploy
npm run deployThis applies remote migrations, builds the Worker, and deploys. The first deploy also creates the Durable Object namespaces and rate-limit bindings declared in wrangler.jsonc.
8. Bootstrap the first user
bland authenticates through tessera. After the first deploy, sign in through tessera with a verified account:
- Confirm tessera has the bland redirect URI registered.
- Open
https://bland.tools/login. - Complete tessera sign-in.
bland will create the user, default workspace, and owner membership from the valid tessera sub. Then issue invites to everyone else from workspace settings.
9. Post-setup smoke check
Run the Post-Deploy Smoke Checks below to confirm auth, collaboration, sharing, uploads, and search are all live.
Deploy Runbook
There is no separate staging Wrangler environment configured in the live repo today. The default config is the production target.
Preflight
Run these from repo root before deployment:
npm ci --ignore-scripts
npm run typecheck
npm test
npm run buildStandard production deploy
npm run deployCurrent behavior:
npm run deployrunsnpm run db:migrate:remote- then
npm run build - then
wrangler deploy
Manual split deploy
Use this if you want the migration and deploy steps separated:
npm run db:migrate:remote
npm run build
npx wrangler deployLocal-only migration
For local development only:
npm run db:migrate:localDo not use db:migrate:local as part of production deployment.
Post-Deploy Smoke Checks
Run these checks immediately after production deploy:
GET https://bland.tools/api/v1/health- Load
https://bland.tools/loginand confirm the tessera sign-in action starts/api/v1/oidc/start. - Complete tessera sign-in with a verified account and confirm session refresh works after reload.
- Open a page, edit content, reload, and confirm content persists.
- Open the same page in a second client and confirm collaboration still connects.
- Create or open a share link and verify the expected view/edit behavior.
- Upload an allowed file and confirm the resulting
/uploads/:idURL serves correctly for an authorized viewer. - Edit page text, wait for queue-driven indexing, and verify the page appears in workspace search.
- On a page with content, select text and run an AI rewrite (e.g. Proofread) from the formatting toolbar; confirm the suggestion streams in and accept/reject works.
- Trigger a slash-menu AI generation (e.g.
/continue) and confirm streaming insertion at the cursor. - Open the page summarize sheet and confirm a summary streams in. Ask one follow-up question and confirm the answer streams.
- Confirm AI actions are not exposed on the share-link surface (
/s/:token).
Important:
/api/v1/healthis only a liveness check. It currently returnsstatusandtimestamponly.- A passing health check does not prove D1, R2, Queue, or Durable Object correctness.
Observability
Primary tools
- Worker logs:
npx wrangler tail bland --format pretty- Deployment status:
npx wrangler deployments status
npx wrangler deployments list- Queue metadata:
npx wrangler queues info bland-tasksBackend signals to watch first
High-signal Worker and DO events in the live source tree:
unhandled_errormessage_failedqueue_send_failedsnapshot_save_failedtitle_sync_failedindex_page_failedsearch_failedworkspace_indexer_clear_failedrate_limit_exceededorigin_rejectedauth_failedoidc_start,oidc_callback_successdiscovery_failed,oidc_token_exchange_failedaccess_deniedshare_access_deniedai_request,ai_response— every AI call emits a paired request/response log with action, surface, page access, duration, and outcomeai_denied— entitlement gate refused an AI call (e.g. share-surface user attempting rewrite)summarize_failed,ai_chat_failed,ai_chat_stream_failed,ai_client_misconfigured— backend-side AI failures
Client-side signals
Client-side Sentry is enabled through SPA shell bootstrap when SENTRY_DSN is set.
High-signal client capture sources:
client-config.bootstrapwindow.errorwindow.unhandledrejectionreact.root-uncaughtpage.loadshared-page.resolveshared-page.active-page-load
Incident Runbooks
Roll back the Worker deployment
List recent deployments:
npx wrangler deployments listRollback:
npx wrangler rollback <version-id> --name bland --message "rollback reason"This rolls back Worker code, not D1 data, DO-local SQLite state, or R2 objects.
Restore D1 metadata with Time Travel
Inspect the recovery point:
npx wrangler d1 time-travel info bland-prodRestore to a timestamp:
npx wrangler d1 time-travel restore bland-prod --timestamp <RFC3339 timestamp>Notes:
- This acts on the remote D1 database.
- D1 restore only covers relational metadata stored in D1.
- It does not restore
DocSyncdocument snapshots orWorkspaceIndexerFTS data, because those live in Durable Object local SQLite.
Search is stale or missing results
Current state:
- Search is a derived projection owned by
WorkspaceIndexer. - There is no bulk rebuild script checked into this repo today.
- There is no DLQ configured for
bland-tasks.
What you can do now:
- Inspect queue metadata with
npx wrangler queues info bland-tasks. - Tail logs and look for
message_failed,fts_indexed,fts_removed,index_page_failed, andsearch_failed. - Re-save an affected page to enqueue a fresh
index-pagemessage. - Archive and restore an affected page to force removal/re-index behavior if appropriate.
What you cannot do today:
- Run a repo-supported full workspace search rebuild.
Document content incident
Current state:
- Persisted document content lives in
DocSyncDO-local SQLite, not in D1. - There is no repo-local bulk export, bulk restore, or replay script for DocSync state.
Operational guidance:
- Treat this as a Durable Object storage incident, not a D1 incident.
- Use Cloudflare Durable Object SQLite recovery tooling/support outside this repo as applicable.
- Do not assume a D1 restore repairs document content.
Upload or R2 incident
Current state:
- R2 has no time-travel restore path in this repo.
- Missing R2 objects are gone unless recovered outside normal repo tooling.
- The D1
uploadsrow can remain as an audit trail even if the blob is missing.
Useful object commands:
npx wrangler r2 object get <bucket>/<key>
npx wrangler r2 object delete <bucket>/<key>AI request incident
Current state:
- AI is a Worker-side feature backed by the Workers AI binding by default. There is no D1 or DO state for AI; failures are request-scoped.
- Per-user rate limit is
RL_AIat 30/min. - AI is member-only. Share-link surfaces never see AI affordances and the routes refuse them with
ai_denied.
Triage steps:
- Tail logs and look for
ai_responseoutcome=errorpaired with the relevanterrorCode(ai_chat_failed,ai_chat_stream_failed,ai_summarize_failed,ai_misconfigured,ai_backend_failed,rate_limited,page_empty). - If
ai_client_misconfiguredappears, confirmBLAND_AI_MODEisworkers-aiin production and that theAIbinding is present inwrangler.jsonc. - If many users hit
rate_limited, review whetherRL_AI(namespace_id101003, 30/min) is still appropriate or whether a single user is hot. - To temporarily disable all AI features in production without a redeploy, push
BLAND_AI_MODE=mockis not acceptable in production (the mock backend is for E2E only). The current escape hatch is to roll the deployment back to a build before AI features.
Queue delivery incident
Current state:
- Queue failures only affect derived search indexing.
- Source-of-truth user/workspace/page metadata and document snapshots are stored elsewhere.
Useful commands:
npx wrangler queues info bland-tasks
npx wrangler queues pause-delivery bland-tasks
npx wrangler queues resume-delivery bland-tasksBe cautious with purge; it drops pending derived work.
Known Gaps
These are current operational limitations, not bugs in this document:
- No worker-side Sentry integration.
- No staging Wrangler environment in the repo.
/api/v1/healthdoes not verify D1, R2, Queue, or DO dependencies.- No DLQ configured for
bland-tasks. - No repo-supported bulk search rebuild script.
- No repo-supported bulk restore/export path for
DocSyncDO-local document storage. - Workspace deletion clears D1 rows and
WorkspaceIndexer, but does not synchronously delete R2 objects or DocSync DO storage. - Uploads are Worker-proxied through
PUT /uploads/:id/data, not direct presigned browser-to-R2 uploads. - PWA installability/offline shell work is separate from v1 operator responsibilities.
- AI first wave (rewrite, generate, summarize, ask-page) is shipped. Second-wave items remain deferred: semantic search / ask-workspace (no Vectorize binding configured), reusable prompt presets, persistent writing instructions, and edge inference. AI rewrite/generate output is text-parsed (paragraphs and bullet lists only); see
AGENTS.md"Deferred Work" for details.