Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .github/workflows/build_worker.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ on:
- "Cargo.toml"
- ".github/actions/build-worker/action.yml"
- ".github/workflows/build_worker.yml"
- "travel_times/**"
push:
# dev / master は deploy_staging.yml / deploy_production.yml が同じ
# composite action で検証してからデプロイするため、ここでは走らせない。
Expand All @@ -42,6 +43,7 @@ on:
- "Cargo.toml"
- ".github/actions/build-worker/action.yml"
- ".github/workflows/build_worker.yml"
- "travel_times/**"
workflow_dispatch:

name: Build Cloudflare Worker
Expand Down Expand Up @@ -70,6 +72,13 @@ jobs:

- uses: ./.github/actions/build-worker

# 到着時間推定の所要時間の見張り (src/travel_times.rs)。記録
# (travel_times/baseline.csv) は本番と同じ生成データで作るので、上の
# action が generated/ を作ったこのジョブで比べる。バスのフィードが
# 欠けても鉄道の推定は変わらない。
- name: Check travel times against the benchmark
run: cargo test -p stationapi-worker travel_times

- uses: actions/upload-artifact@v4
with:
name: worker-build
Expand Down
1 change: 1 addition & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,7 @@ The Worker is the workspace root package. `stationapi`, `preprocessor`, and `dat
- **Schema** – Changing a GraphQL type changes the SDL. Update `schema/public.graphql` in the same change; CI compares it against the running Worker's `/__schema` and fails on any difference. That diff is exactly the client-visible impact.
- **Data verification** – Execute `cargo run -p data_validator` whenever CSVs change and record results in pull requests.
- **IPA coverage audit** – Execute `make ipa-audit` when English or romanized CSV names change. This is a read-only report for `data/2!lines.csv`, `data/3!stations.csv`, and `data/4!types.csv`; it does not fail validation, but highlights unresolved tokens and example names so the IPA dictionary can be extended deliberately.
- **Travel-time benchmark** – `travel_times/cases.csv` lists real travel times (a range and a typical value — the median — in minutes, weekday daytime, trains with the same stopping pattern as the line group) that the arrival estimation is measured against. `cargo test -p stationapi-worker` (`src/travel_times.rs`) fails when any case moves more than 1 percentage point further from its typical value than `travel_times/baseline.csv` records (a range alone hides drifts inside a range widened by one outlier train), when the mean gets worse, or when a recorded case can no longer be estimated. It compares only when the Worker embeds `generated/` (the estimation uses track lengths and line groups that exist only there), so `build_worker.yml` runs it after building the data; on `data/*.csv` it just prints the table. `make travel-time-report` (`TRAVEL_TIME_API`, default the local `make dev` Worker) measures every case against a Worker built from the generated data. A change to the speed tables, the general speed rules, or the estimation parameters must attach the report from before and after, and must not trade one line's accuracy for the whole set's; after an intended change, run `make data` and regenerate the baseline with `TRAVEL_TIMES_UPDATE_BASELINE=1 cargo test -p stationapi-worker travel_times`. Values come from open GTFS feeds (credit them in the README's Data Sources) or from the maintainer.
- **Endpoint benchmarks** – `make bench` (or `python3 .claude/skills/benchmark-gql/bench.py`) replays every `Query` field against production (`gql.trainlcd.app`, script `stationapi`) and staging (`gql-stg.trainlcd.app`, script `stationapi-stg`) and writes a Markdown report under `benchmarks/`. Both environments embed the same data, so any difference is implementation — which makes this the way to see what a `dev`-to-`master` release will do to performance before it ships. Besides client latency it records the Worker's `cpuTime`, read from `wrangler tail --format json` and matched to each request by `cf-ray`; the tail is filtered on a per-run request header, so production's live traffic does not leak into the sample. Collecting CPU time needs the `workers_tail (read)` scope, and the run sends hundreds of real requests to production — it is not a routine check. Add a case to `.claude/skills/benchmark-gql/queries.json` whenever a `Query` field is added, and never edit an existing case's variables: the reports are meant to stay comparable across runs.

## GraphQL Query Overview
Expand Down
10 changes: 9 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# StationAPI Makefile
# よく使うタスクの定義

.PHONY: help test check fmt clippy data build dev deploy deploy-production schema ipa-audit bench clean
.PHONY: help test check fmt clippy data build dev deploy deploy-production schema ipa-audit bench travel-time-report clean

# CI (.github/workflows/build_worker.yml) と同じ版を使う。グローバルへ入れて
# いなくても npx が取ってくるので、版ずれでビルド結果が変わらない。
Expand All @@ -22,6 +22,7 @@ help:
@echo " schema - Diff the running Worker's SDL against schema/public.graphql"
@echo " ipa-audit - Print IPA coverage report for English/romanized CSV names"
@echo " bench - Compare production vs staging GraphQL performance (sends live traffic to both)"
@echo " travel-time-report - Compare estimated travel times with travel_times/cases.csv (TRAVEL_TIME_API, default http://127.0.0.1:8787/)"
@echo " clean - Clean build artifacts"
@echo ""
@echo "Environment variables:"
Expand Down Expand Up @@ -84,6 +85,13 @@ ipa-audit:
# 実在のエンドポイントへ数百リクエスト投げるので、気軽に回すものではない。
# CPU Time の収集には wrangler の workers_tail (read) 権限が要る。
# 追加の引数は BENCH_ARGS で渡す (例: make bench BENCH_ARGS="--repeat 30")。
# 到着時間推定の所要時間を、実際の所要時間 (travel_times/cases.csv) と比べる。
# 本番と同じ生成データで測るため、`make data && make dev` で起動した Worker か
# ステージングに向ける (TRAVEL_TIME_API)。
TRAVEL_TIME_API ?= http://127.0.0.1:8787/
travel-time-report:
python3 scripts/travel_time_report.py --api "$(TRAVEL_TIME_API)"

bench:
@echo "警告: 本番 (gql.trainlcd.app) とステージングへ実リクエストを送ります。" >&2
@echo " 既定で 1 環境あたり 400 件超、うち数十件は Worker の CPU を 500ms 以上使います。" >&2
Expand Down
5 changes: 5 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,11 @@ This project includes a comprehensive dataset of Japanese railway information in
(https://nlftp.mlit.go.jp/ksj/gml/datalist/KsjTmplt-N02-2025.html) を加工して作成
- Bus stops and routes are derived from the GTFS and ODPT feeds listed in
`preprocessor/src/gtfs/feed.rs` and `preprocessor/src/gtfs/odpt.rs`.
- Some of the reference travel times in `travel_times/cases.csv` are derived from
the Toei Subway GTFS (Bureau of Transportation, Tokyo Metropolitan Government,
CC BY 4.0) and from the Tokyo Metro and Metropolitan Intercity Railway
(Tsukuba Express) GTFS feeds published by the Public Transportation Open Data
Center under the Public Transportation Open Data Basic License.

## Contributors ✨

Expand Down
10 changes: 10 additions & 0 deletions build.rs
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,16 @@ fn main() {
staged.len()
);
}
// どちらのデータを埋め込んだか。所要時間のベンチマーク (src/travel_times.rs) の
// 記録は生成データで作るので、data/*.csv のときは記録と比べない
println!(
"cargo:rustc-env=STATIONAPI_EMBEDDED_DATA={}",
if generated_count == 0 {
"data"
} else {
"generated"
}
);
if generated_count == 0 {
println!(
"cargo:warning=generated が無いため data/*.csv を使用します。\
Expand Down
17 changes: 17 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -506,6 +506,23 @@ input RouteLegInput { lineGroupId: Int! fromStationId: Int! toStationId: Int!
- 両端の駅が `fromStationId` / `toStationId` と一致しない
- `viaLineIds`・`directionId`・`lineGroupId` と同時に指定されている

### 所要時間のベンチマーク (`travel_times/`)

到着時間推定の所要時間を、実際の列車の所要時間と比べる基準を `travel_times/cases.csv`
に置いています。速度の較正テーブルや一般則は、1 つの路線に合わせて変えると、同じ
規則を使うほかの路線の推定も変わります。変更の前後で全体の誤差を測るための仕組み
です。

- `cargo test -p stationapi-worker` (`src/travel_times.rs`) は、基準ごとの「実際の
典型的な所要時間からのずれ」を `travel_times/baseline.csv` の記録と比べ、悪くなる
と失敗します。記録は本番と同じ生成データで作るので、比べるのは `generated/` で
動くときだけです。
CI では `build_worker.yml` が生成データを作ってから走らせます。
- `make travel-time-report` は、生成データで動く Worker に問い合わせて全件の誤差を
出します。推定の規則や較正を変える PR には、変更前と変更後のレポートを載せます。

基準の決め方と記録の更新方法は `travel_times/README.md` にあります。

### 行き先の検索 (`stationsByName`)

`stationsByName` に `fromStationGroupId` を指定すると、その駅から行ける駅だけに
Expand Down
109 changes: 109 additions & 0 deletions scripts/travel_time_report.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
#!/usr/bin/env python3
"""実際の所要時間 (travel_times/cases.csv) に対する到着時間推定の誤差を、動いている
Worker に問い合わせて Markdown で出す。

CI の回帰テスト (src/travel_times.rs) は data/*.csv だけで動くので、生成データにしか
無い種別グループを飛ばし、線路の長さも持たない。本番と同じ生成データでの精度は、
`make data && make dev` で起動した Worker か、ステージングに向けてこれで測る。
推定の規則や較正を変える PR には、変更前と変更後のこのレポートを載せる。

使い方:
python3 scripts/travel_time_report.py # http://127.0.0.1:8787/
python3 scripts/travel_time_report.py --api <URL>
"""
from __future__ import annotations

import argparse
import csv
import json
import statistics
import sys
import urllib.request
from pathlib import Path

CASES = Path(__file__).resolve().parent.parent / "travel_times" / "cases.csv"
# 既定の Python-urllib は配信側で弾かれるので、bench.py と同じく名乗る
USER_AGENT = "stationapi-travel-time-report/1.0 (+https://github.com/TrainLCD/StationAPI)"
QUERY = """query TravelTimeReport($from: Int!, $to: Int!, $group: Int!) {
estimateArrivalTimes(fromStationId: $from, toStationId: $to,
legs: [{ lineGroupId: $group, fromStationId: $from, toStationId: $to }]) {
routes { stops { stationId cumulativeMinutes departureCumulativeMinutes } }
}
}"""


def estimate(api: str, case: dict) -> float:
end = int(case["slice_end_station_id"] or case["to_station_id"])
body = json.dumps({
"query": QUERY,
"variables": {
"from": int(case["from_station_id"]),
"to": end,
"group": int(case["line_group_id"]),
},
}).encode()
req = urllib.request.Request(
api, data=body, headers={"content-type": "application/json", "user-agent": USER_AGENT}
)
with urllib.request.urlopen(req, timeout=30) as res:
data = json.load(res)
if data.get("errors"):
raise RuntimeError("; ".join(e.get("message", "") for e in data["errors"]))
stops = data["data"]["estimateArrivalTimes"]["routes"][0]["stops"]
target = int(case["to_station_id"])
stop = next(s for s in stops[1:] if s["stationId"] == target)
key = "cumulativeMinutes" if case["measure"] == "arrival" else "departureCumulativeMinutes"
return float(stop[key])


def range_error(est: float, lo: float, hi: float) -> float:
if est < lo:
return (lo - est) / lo
if est > hi:
return (est - hi) / hi
return 0.0


def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--api", default="http://127.0.0.1:8787/")
args = parser.parse_args()

with CASES.open(encoding="utf-8") as f:
cases = list(csv.DictReader(f))

rows, range_errors, typical_errors, failed = [], [], [], []
for case in cases:
lo, hi = float(case["real_min_minutes"]), float(case["real_max_minutes"])
typical = float(case["real_typical_minutes"])
try:
est = estimate(args.api, case)
except Exception as e: # noqa: BLE001 - 1 件の失敗で全体を止めない
failed.append(f"{case['label']}: {e}")
continue
t, r = (est - typical) / typical, range_error(est, lo, hi)
typical_errors.append(abs(t))
range_errors.append(r)
rows.append(
f"| {case['label']} | {typical:g}分 ({lo:g}〜{hi:g}分) | {est:.1f}分 "
f"| {t * 100:+.1f}% | {r * 100:.1f}% |"
)

print(f"# 到着時間推定の誤差 ({args.api})\n")
print("| 基準 | 実際の典型 (範囲) | 推定 | 典型からのずれ | 範囲からの外れ |")
print("| --- | --- | --- | --- | --- |")
print("\n".join(rows))
if typical_errors:
print(
f"\n{len(typical_errors)} 件: 典型からのずれ (絶対値) の平均 "
f"{statistics.mean(typical_errors) * 100:.2f}%、"
f"範囲からの外れの平均 {statistics.mean(range_errors) * 100:.2f}%"
)
if failed:
print("\n推定できなかった基準:\n")
print("\n".join(f"- {line}" for line in failed))
return 0 if typical_errors else 1


if __name__ == "__main__":
sys.exit(main())
2 changes: 2 additions & 0 deletions src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,8 @@
mod graphql;
mod index;
mod repository;
#[cfg(test)]
mod travel_times;

use async_graphql::http::GraphiQLSource;
use async_graphql::Request as GqlRequest;
Expand Down
Loading
Loading