Skip to content

feat(adoption): monthly HTTP Archive adoption numbers for spec topics - #243

Merged
jdevalk merged 9 commits into
mainfrom
feat/httparchive-adoption
Sep 30, 2026
Merged

jdevalk merged 9 commits into
mainfrom
feat/httparchive-adoption

Conversation

@jdevalk

@jdevalk jdevalk commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Adds monthly HTTP Archive adoption percentages to spec pages for desktop root pages with rank <= 1000000. The display shows the actual sample size and crawl month, using “websites” in the reader-facing wording.

Query and reporting

  • One BigQuery aggregation over the newest available httparchive.crawl.pages partition writes src/data/adoption.json.
  • Nine topics use only custom_metrics.well_known and custom_metrics.robots_txt. llms.txt is excluded to avoid reading the large custom_metrics.other column.
  • Both numerator and denominator use the same filters. Distinct root_page values retain scheme and port.
  • ADOPTION_DRY_RUN=1 validates without executing or writing data; ADOPTION_PRINT_QUERY=1 prints SQL offline; ADOPTION_CRAWL=YYYY-MM selects a month. Execution defaults to the authorised 112 GiB cap, which accommodates the clustered-table planning estimate.

Data quality

  • HTTP 200 alone no longer counts as adoption. Security.txt requires the collector’s parsed field checks; WebAuthn requires a non-empty origins array; change-password requires a followed redirect and a 404 from the collector’s deliberately nonexistent URL.
  • AI Catalog and ARD manifests both belong to Agentic Resource Discovery. Either needs at least one parsed entry; an origin with both counts once. They never contribute to the unrelated RFC 9727 API Catalog topic.
  • GPC and app-association documents require meaningful parsed fields. Robots.txt requires parsed directives, and AI crawler groups require a successful response.
  • API Catalog and the six topics awaiting the upstream collector extension are omitted instead of represented as zero adoption. Future metrics need reviewed detection rules before inclusion.
  • These are conservative detection signals, not conformance certification. Empty manifests, empty robots files, missing parser fields and inconclusive probes may undercount implementations. Exact criteria and pinned collector sources are documented in scripts/adoption/README.md.

Scan size

Reviewer feedback reports 391.66 GB for the original query and 1.52 GB when excluding llms.txt and restricting rank. The earlier 1.93 TiB claim was incorrect.

The revised September 2026 query passes a free dry run returning 119,302,837,357 bytes, explicitly marked UPPER_BOUND. The table is clustered by client, root-page status, rank and page. This upper bound is not actual bytes scanned. The initial execution on 30 September 2026 failed at the default 100 GiB cap because BigQuery required an upper-bound limit of 119,303,831,552 bytes. An explicitly authorised one-off retry with a 112 GiB cap succeeded without a cache hit. Job openklauw:US.adoption_verify_112gib_20260930_1790795002586 processed 1,634,231,969 bytes (1.522 GiB) and billed 1,634,729,984 bytes. This confirms the roughly 1.52 GiB actual scan despite the much larger planning estimate.

The result covers 668,524 desktop root pages and 668,523 distinct origins in the September 2026 top-million-ranked sample. All counts and rounded percentages were validated against their denominators.

Signal Detected origins Percentage
gpc-json 40,017 5.986%
security-txt 17,997 2.692%
assetlinks-json 37,595 5.624%
apple-app-site-association 81,422 12.179%
change-password 3,910 0.585%
webauthn 345 0.052%
agentic-resource-discovery 0 0%
robots-txt 499,463 74.711%
robots-for-ai-crawlers 87,303 13.059%

ARD's zero is zero detected evidence under these rules, not proof of no adoption: older crawls may lack its newer parsed fields. The verified September report is now committed in src/data/adoption.json. The authorised default monthly cap is 112 GiB (120,259,084,288 bytes). Eight topic pages display adoption percentages beneath their summaries, including the month and actual desktop sample size; zero or unmeasured results are hidden. The PR has not been merged.

Monthly workflow

Runs on the 18th at 06:00 UTC, supports manual runs and opens or updates a data-refresh PR. GCP authentication uses Workload Identity Federation; required repository variables are documented in the README.

Validation

  • All 11 tests pass with the optional BigQuery integration enabled. Twenty fixtures reproduce upstream collector output for HTML and JSON catch-alls, redirects, valid documents, failed probes and both ARD paths.

  • The integration executes generated SQL using only inline parameters after asserting a zero-byte dry run. It verifies false-positive rejection, distinct-origin counts, ARD deduplication, sample boundaries and percentages without reading HTTP Archive.

  • CI runs the 10 offline tests; the authenticated integration is opt-in with ADOPTION_TEST_PROJECT.

  • Live HTTP Archive query dry run, ESLint, Prettier, skill integrity, Astro check (0 errors) and complete site build pass locally.

  • Populated build verified: all eight displayed percentages match the real report; ARD and unmeasured API Catalog show no adoption line. Browser preview confirms the security.txt figure (2.692%) beneath its summary.

Adds a monthly job that queries the HTTP Archive's custom metrics in
BigQuery and refreshes src/data/adoption.json, which the spec pages
render as an "Adoption: X% of origins" line under the summary.

- scripts/adoption/metrics.json: spec slug -> BigQuery JSONPath conditions
  (17 topics: well-known paths, llms.txt, robots.txt, AI-crawler rules)
- scripts/adoption/fetch-adoption.mjs: single aggregation query over the
  newest httparchive.pages crawl table; npm run adoption
- .github/workflows/adoption-monthly.yml: runs on the 18th, opens/updates
  a PR like the Plausible refresh job; auth via Workload Identity
  Federation (setup in scripts/adoption/README.md)
- SpecLayout renders the line when data exists for the page's slug

Query cost is one small aggregation per month, inside BigQuery's free tier.
@jdevalk

jdevalk commented Sep 28, 2026

Copy link
Copy Markdown
Owner Author

The monthly workflow file that could not be pushed (missing workflow scope on this machine's token). To enable the automation: on GitHub, open this branch, Add file → Create new file, name it .github/workflows/adoption-monthly.yml, paste the content below, and commit to this branch.

name: Adoption refresh

# Pulls HTTP Archive custom-metric adoption numbers for spec topics
# (scripts/adoption/) and opens (or updates) a single PR refreshing
# src/data/adoption.json, which the spec pages render automatically.
#
# Runs monthly on the 18th: the crawl for month M lands in BigQuery around
# mid-month M+1, so the 18th reliably picks up the newest complete crawl.
# The fetch script itself also falls back to the newest table it can find.
#
# One-time setup is documented in scripts/adoption/README.md (GCP project +
# Workload Identity Federation; no long-lived secrets).

on:
  schedule:
    - cron: "0 6 18 * *" # 18th of each month, 06:00 UTC
  workflow_dispatch:

permissions:
  contents: write
  pull-requests: write
  id-token: write

jobs:
  refresh:
    name: Refresh adoption data
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7

      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm

      - uses: google-github-actions/auth@v2
        with:
          project_id: ${{ vars.GCP_PROJECT_ID }}
          workload_identity_provider: ${{ vars.GCP_WIF_PROVIDER }}
          service_account: ${{ vars.GCP_SERVICE_ACCOUNT }}

      - name: Install dependencies
        run: npm ci

      - name: Fetch adoption data
        run: npm run adoption
        env:
          GCP_PROJECT_ID: ${{ vars.GCP_PROJECT_ID }}

      - name: Format data file
        run: npx prettier --write src/data/adoption.json

      - name: Detect change
        id: diff
        run: |
          if git diff --quiet -- src/data/adoption.json; then
            echo "changed=false" >> "$GITHUB_OUTPUT"
            echo "Adoption data unchanged."
          else
            echo "changed=true" >> "$GITHUB_OUTPUT"
          fi

      - name: Open or update PR
        if: steps.diff.outputs.changed == 'true'
        env:
          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
        run: |
          branch="chore/refresh-adoption"
          crawl=$(node -p "JSON.parse(require('fs').readFileSync('src/data/adoption.json','utf8')).crawl")
          git config user.name "github-actions[bot]"
          git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
          git checkout -B "$branch"
          git add src/data/adoption.json
          git commit -m "chore(adoption): refresh HTTP Archive adoption data ($crawl)"
          git push -f -u origin "$branch"
          body="Monthly refresh of the HTTP Archive adoption numbers rendered on the spec pages (\`src/data/adoption.json\`, crawl \`$crawl\`).

          Generated by \`npm run adoption\` — a single BigQuery aggregation over the newest \`httparchive.pages\` crawl table. Merge to deploy. CI verifies the build."
          gh pr view "$branch" >/dev/null 2>&1 \
            && gh pr edit "$branch" --body "$body" \
            || gh pr create --base main --head "$branch" \
                 --title "chore(adoption): refresh HTTP Archive adoption data ($crawl)" \
                 --body "$body"

Then follow the one-time GCP setup in scripts/adoption/README.md (project + Workload Identity Federation + three repo variables). The next scheduled run (18th of the month, or workflow_dispatch) will open the first data PR.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Deploying specification-website with  Cloudflare Pages  Cloudflare Pages

Latest commit: a56f455
Status: ✅  Deploy successful!
Preview URL: https://4c97d132.specification-website.pages.dev
Branch Preview URL: https://feat-httparchive-adoption.specification-website.pages.dev

View logs

@jdevalk
jdevalk merged commit b669a14 into main Sep 30, 2026
9 checks passed
@jdevalk
jdevalk deleted the feat/httparchive-adoption branch September 30, 2026 19:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant