Skip to content

ci: fail the EIF build when the SSH or Docker daemon wait times out - #82

Open
nickpell wants to merge 1 commit into
mainfrom
nick/eif-wait-ssh-fail
Open

nickpell wants to merge 1 commit into
mainfrom
nick/eif-wait-ssh-fail

Conversation

@nickpell

@nickpell nickpell commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

  • The "Wait for SSH" step in .github/workflows/eif-build.yml tried SSH 30 times and then passed even when every attempt failed, so the job failed one step later in "Copy scripts to instance" with an scp connection error. The step now records a successful attempt and, when all 30 attempts fail, prints ::error::SSH to the EIF builder instance did not succeed after 30 attempts and exits 1. A successful attempt still breaks out of the loop at once.
  • install_docker in enclave/scripts/setup-nitro-instance.sh has two Docker daemon waits with the same fall-through: one after it starts an installed but stopped Docker service, and one after a fresh install. After 10 failed docker info probes the script carried on, verify_setup checks only command -v docker, and the first daemon error appeared in the "Build EIF" step. Each wait now logs ERROR: Docker daemon did not respond to 'docker info' after 10 attempts and returns 1. main calls install_docker as a plain command under set -euo pipefail, so the script exits 1 there, and the "Setup Nitro instance" step fails with it because ssh returns the remote exit status.
  • Attempt counts and delays are unchanged. The cleanup steps for the instance, security group and key pair run with if: always(), so they still run when either step fails. The script runs only on the ephemeral builder instance and is not an EIF input (enclave/Dockerfile copies only the enclave binary), so the PCRs do not change.
  • The same change for openarbiter is in ci: fail the EIF build when the SSH or Docker daemon wait times out openarbiter#39. The two PRs are independent and can merge in either order.

Notes for reviewers

Pre-merge checklist

  • Lint passes: actionlint .github/workflows/eif-build.yml (which runs shellcheck on run: blocks) reports the same 2 pre-existing notes on main and on this branch, none in "Wait for SSH", and exits 0 with -shellcheck=. shellcheck enclave/scripts/setup-nitro-instance.sh reports only the existing SC1091 note for /etc/os-release. The added lines are already in shfmt -i 4 -sr form. mise run //:ratchet:lint passes.
  • "Wait for SSH" behaviour: the step's run: block, extracted from the workflow and run under bash -e (the step's shell) with a stub ssh, exits 0 after 1 call when SSH answers at once, exits 0 after 3 calls when SSH answers on attempt 3, and prints the ::error:: line and exits 1 after 30 calls when SSH never answers. The block on main exits 0 in that last case.
  • Docker wait behaviour: the whole script, run with stub sudo, systemctl, dnf and docker, completes with exit 0 when docker info succeeds on attempt 2 or 3, for both waits. When docker info never succeeds, it logs the error after 10 probes and exits 1 before install_dependencies, for both waits. The script on main logs ✓ Setup complete! and exits 0 in that case.
  • Diff contains no unintended changes: two files, 18 added lines and no removed lines, identical to ci: fail the EIF build when the SSH or Docker daemon wait times out openarbiter#39 apart from the script path. git diff --check passes.
  • PR checks pass on head 48f7570: Go (test, lint) and Ratchet.

Post-deploy/apply verification

The EIF workflow runs only from main (workflow_run after Docker Build, or workflow_dispatch), so this PR's checks do not run the edited steps.

  • The next Build EIF run on main after the merge is green, including update-pcrs, and its log shows SSH is ready! in "Wait for SSH" and ✓ Docker daemon started (or ✓ Docker service already running) in "Setup Nitro instance".

The Wait for SSH step and the two Docker daemon waits in setup-nitro-instance.sh fell through after their last attempt, so the job failed in a later step. Each loop now records success and fails with an error when every attempt has failed.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The timeout handling is correct, scoped, and preserves existing retry behavior.

Review effort: Balanced
Findings: None

What changed in this PR

Ensures EIF CI fails immediately when SSH or Docker readiness probes time out.

Changes:

  • Track successful SSH and Docker probes.
  • Exit with explicit errors after exhausting retries.
File Description
.github/​workflows/​eif-build.yml Fails the job when SSH never becomes ready.
enclave/​scripts/​setup-nitro-instance.sh Fails setup when Docker remains unavailable.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants