Skip to content

docs(deploy): add sandbox deploy runbook and recreate-relay script - #768

Open
Ferryx349 wants to merge 3 commits into
mainfrom
feat/deploy-runbook
Open

Ferryx349 wants to merge 3 commits into
mainfrom
feat/deploy-runbook

Conversation

@Ferryx349

Copy link
Copy Markdown
Collaborator

Description

Document the manual post-pull migrate and recreate flow, rollback steps, and /readyz verification. Closes the gap where the webhook pull stack loads images but does not restart the relay.

Related Issue

closes:- #767

Motivation and Context

How Has This Been Tested?

Screenshots (if appropriate):

Types of changes

  • Non-functional change (docs, style, minor refactor)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)

Checklist:

  • My code follows the code style of this project.
  • My change requires a change to the documentation.
  • I have updated the documentation accordingly.
  • I have read the CONTRIBUTING document.
  • I have added tests to cover my code changes.
  • I added a changeset, or this is docs-only and I added an empty changeset.
  • All new and existing tests passed.

@changeset-bot

changeset-bot Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 622e2ba

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 1 package
Name Type
nostream Minor

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@Ferryx349
Ferryx349 force-pushed the feat/deploy-runbook branch from da7ac54 to df40b32 Compare October 4, 2026 05:02
Document post-image update steps for single-relay and HAProxy stacks;
link deploy README to the runbook without duplicating health/HAProxy setup.
@Ferryx349
Ferryx349 force-pushed the feat/deploy-runbook branch from df40b32 to 4baf9b9 Compare October 4, 2026 05:10
@greptile-apps

greptile-apps Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Low impact] Adds deployment documentation and a helper script.

Fix the HAProxy rollback sequence before merging; it cannot restore a stopped green backend as written.

Findings

  1. P1 Rollback cannot restore stopped green ▶
  2. P2 HAProxy files stay outdated ▶
  3. P2 Readiness wait can hang ▶

Summary

Adds a manual deployment runbook and a single-relay recreation script.

  • The single-relay script migrates, recreates, and checks the relay.
  • Deploy guides now explain how to update and roll back relay hosts.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Load new image] --> B[Run migrations]
  B -->|Fail| C[Stop deployment]
  B -->|Pass| D[Recreate relay]
  D --> E[Check readyz]
  E -->|Ready| F[Deployment complete]
  E -->|Not ready| G[Restore previous image]
  G --> H[Recreate without migrations]
Loading

Reviews (2) · Last reviewed commit: "Merge branch 'main' into feat/deploy-run..." · Reviewed by Greptile

Comment thread docs/DEPLOY-RUNBOOK.md Outdated
Comment thread docs/DEPLOY-RUNBOOK.md Outdated
Comment thread deploy/recreate-relay.sh Outdated
Comment thread docs/DEPLOY-RUNBOOK.md Outdated
Capture the running relay image before deploy, document script install path,
use compose exit-code-from for migrate, and add SKIP_MIGRATE for rollback.
Comment thread docs/DEPLOY-RUNBOOK.md
Comment on lines +178 to +180
HAProxy: run migrate only when moving **forward** on a new image. For rollback,
retag the old image to `:main`, then replace relays with
`rolling-relay-recreate.sh` (it does not re-run `nostream-migrate`).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Rollback cannot restore stopped green

If nostream-green fails after replacement, these instructions tell the operator to stop it and run rolling-relay-recreate.sh. The script starts with the still-running nostream-blue and exits because its peer is stopped. Neither backend gets restored.

Document how to recreate the stopped backend on the old image first, then replace the other backend.

Comment thread docs/DEPLOY-RUNBOOK.md
Comment on lines +62 to +68
## Shared steps (both stacks)

### 1. Refresh release-managed files (when release notes say so)

```bash
./deploy/bootstrap.sh /opt/nostream
```

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 HAProxy files stay outdated

The shared refresh step tells both stacks to run bootstrap.sh, but that script only copies the single-relay Compose file and postgresql.conf. It does not refresh docker-compose.haproxy.yml, rolling-relay-recreate.sh, or the HAProxy files.

If a release changes those files, an HAProxy host keeps the old copies despite following this step. Add separate HAProxy refresh commands so operators do not have to discover the missing steps themselves.

Comment thread deploy/recreate-relay.sh
Comment on lines +41 to +43
if curl -sf "$READYZ_URL" >/dev/null; then
echo "ready: $READYZ_URL"
curl -s "$READYZ_URL"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Readiness wait can hang

Neither readiness request has a timeout. If the relay accepts a connection but never finishes responding, curl can wait indefinitely. READYZ_RETRIES cannot advance the loop, and the operator never gets the final failure diagnostics.

Add a request timeout to both readiness requests.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant