Repository navigation
feat(adoption): monthly HTTP Archive adoption numbers for spec topics - #243
Merged
Merged
Conversation
Adds a monthly job that queries the HTTP Archive's custom metrics in BigQuery and refreshes src/data/adoption.json, which the spec pages render as an "Adoption: X% of origins" line under the summary. - scripts/adoption/metrics.json: spec slug -> BigQuery JSONPath conditions (17 topics: well-known paths, llms.txt, robots.txt, AI-crawler rules) - scripts/adoption/fetch-adoption.mjs: single aggregation query over the newest httparchive.pages crawl table; npm run adoption - .github/workflows/adoption-monthly.yml: runs on the 18th, opens/updates a PR like the Plausible refresh job; auth via Workload Identity Federation (setup in scripts/adoption/README.md) - SpecLayout renders the line when data exists for the page's slug Query cost is one small aggregation per month, inside BigQuery's free tier.
Owner
Author
|
The monthly workflow file that could not be pushed (missing name: Adoption refresh
# Pulls HTTP Archive custom-metric adoption numbers for spec topics
# (scripts/adoption/) and opens (or updates) a single PR refreshing
# src/data/adoption.json, which the spec pages render automatically.
#
# Runs monthly on the 18th: the crawl for month M lands in BigQuery around
# mid-month M+1, so the 18th reliably picks up the newest complete crawl.
# The fetch script itself also falls back to the newest table it can find.
#
# One-time setup is documented in scripts/adoption/README.md (GCP project +
# Workload Identity Federation; no long-lived secrets).
on:
schedule:
- cron: "0 6 18 * *" # 18th of each month, 06:00 UTC
workflow_dispatch:
permissions:
contents: write
pull-requests: write
id-token: write
jobs:
refresh:
name: Refresh adoption data
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm
- uses: google-github-actions/auth@v2
with:
project_id: ${{ vars.GCP_PROJECT_ID }}
workload_identity_provider: ${{ vars.GCP_WIF_PROVIDER }}
service_account: ${{ vars.GCP_SERVICE_ACCOUNT }}
- name: Install dependencies
run: npm ci
- name: Fetch adoption data
run: npm run adoption
env:
GCP_PROJECT_ID: ${{ vars.GCP_PROJECT_ID }}
- name: Format data file
run: npx prettier --write src/data/adoption.json
- name: Detect change
id: diff
run: |
if git diff --quiet -- src/data/adoption.json; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "Adoption data unchanged."
else
echo "changed=true" >> "$GITHUB_OUTPUT"
fi
- name: Open or update PR
if: steps.diff.outputs.changed == 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
branch="chore/refresh-adoption"
crawl=$(node -p "JSON.parse(require('fs').readFileSync('src/data/adoption.json','utf8')).crawl")
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git checkout -B "$branch"
git add src/data/adoption.json
git commit -m "chore(adoption): refresh HTTP Archive adoption data ($crawl)"
git push -f -u origin "$branch"
body="Monthly refresh of the HTTP Archive adoption numbers rendered on the spec pages (\`src/data/adoption.json\`, crawl \`$crawl\`).
Generated by \`npm run adoption\` — a single BigQuery aggregation over the newest \`httparchive.pages\` crawl table. Merge to deploy. CI verifies the build."
gh pr view "$branch" >/dev/null 2>&1 \
&& gh pr edit "$branch" --body "$body" \
|| gh pr create --base main --head "$branch" \
--title "chore(adoption): refresh HTTP Archive adoption data ($crawl)" \
--body "$body"Then follow the one-time GCP setup in |
Deploying specification-website with
|
| Latest commit: |
a56f455
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://4c97d132.specification-website.pages.dev |
| Branch Preview URL: | https://feat-httparchive-adoption.specification-website.pages.dev |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds monthly HTTP Archive adoption percentages to spec pages for desktop root pages with
rank <= 1000000. The display shows the actual sample size and crawl month, using “websites” in the reader-facing wording.Query and reporting
httparchive.crawl.pagespartition writessrc/data/adoption.json.custom_metrics.well_knownandcustom_metrics.robots_txt.llms.txtis excluded to avoid reading the largecustom_metrics.othercolumn.root_pagevalues retain scheme and port.ADOPTION_DRY_RUN=1validates without executing or writing data;ADOPTION_PRINT_QUERY=1prints SQL offline;ADOPTION_CRAWL=YYYY-MMselects a month. Execution defaults to the authorised 112 GiB cap, which accommodates the clustered-table planning estimate.Data quality
scripts/adoption/README.md.Scan size
Reviewer feedback reports 391.66 GB for the original query and 1.52 GB when excluding
llms.txtand restricting rank. The earlier 1.93 TiB claim was incorrect.The revised September 2026 query passes a free dry run returning 119,302,837,357 bytes, explicitly marked
UPPER_BOUND. The table is clustered by client, root-page status, rank and page. This upper bound is not actual bytes scanned. The initial execution on 30 September 2026 failed at the default 100 GiB cap because BigQuery required an upper-bound limit of 119,303,831,552 bytes. An explicitly authorised one-off retry with a 112 GiB cap succeeded without a cache hit. Jobopenklauw:US.adoption_verify_112gib_20260930_1790795002586processed 1,634,231,969 bytes (1.522 GiB) and billed 1,634,729,984 bytes. This confirms the roughly 1.52 GiB actual scan despite the much larger planning estimate.The result covers 668,524 desktop root pages and 668,523 distinct origins in the September 2026 top-million-ranked sample. All counts and rounded percentages were validated against their denominators.
ARD's zero is zero detected evidence under these rules, not proof of no adoption: older crawls may lack its newer parsed fields. The verified September report is now committed in
src/data/adoption.json. The authorised default monthly cap is 112 GiB (120,259,084,288 bytes). Eight topic pages display adoption percentages beneath their summaries, including the month and actual desktop sample size; zero or unmeasured results are hidden. The PR has not been merged.Monthly workflow
Runs on the 18th at 06:00 UTC, supports manual runs and opens or updates a data-refresh PR. GCP authentication uses Workload Identity Federation; required repository variables are documented in the README.
Validation
All 11 tests pass with the optional BigQuery integration enabled. Twenty fixtures reproduce upstream collector output for HTML and JSON catch-alls, redirects, valid documents, failed probes and both ARD paths.
The integration executes generated SQL using only inline parameters after asserting a zero-byte dry run. It verifies false-positive rejection, distinct-origin counts, ARD deduplication, sample boundaries and percentages without reading HTTP Archive.
CI runs the 10 offline tests; the authenticated integration is opt-in with
ADOPTION_TEST_PROJECT.Live HTTP Archive query dry run, ESLint, Prettier, skill integrity, Astro check (0 errors) and complete site build pass locally.
Populated build verified: all eight displayed percentages match the real report; ARD and unmeasured API Catalog show no adoption line. Browser preview confirms the security.txt figure (2.692%) beneath its summary.