Skip to content

Latest commit

 

History

History
129 lines (112 loc) · 36.4 KB

File metadata and controls

129 lines (112 loc) · 36.4 KB

StationAPI Repository Guidelines

This guide explains how automation agents and human contributors should work with the StationAPI repository so releases stay predictable, auditable, and safe. Update this file whenever you change the workflow or behavior it documents.

Project Layout

  • src/ – The Worker itself (stationapi-worker, wasm32 only). lib.rs holds the endpoints, index.rs parses the embedded CSVs into in-memory indexes, repository.rs implements the repository traits against those indexes, and graphql/ holds the async-graphql types and resolvers. index.rs also holds the spatial grid used by every coordinate lookup — see Coordinate lookups below.
  • schema/public.graphql – The published GraphQL schema. CI diffs the Worker's SDL against this file, so an unintended change fails the build.
  • build.rs – Stages generated/*.csv (falling back to data/*.csv) into OUT_DIR and pre-converts station_station_types and connections into fixed-width binaries.
  • wrangler.jsonc – Staging and production deployment settings.
  • stationapi/src/domain/ – Entity definitions and repository abstractions. repository/ provides async_trait-based interfaces, and normalize.rs contains text normalization for search.
  • stationapi/src/use_case/ – Application logic. interactor/query.rs implements the QueryUseCase contract defined in traits/query.rs; dto/ converts entities into model types (this is where IPA and TTS segments are built).
  • stationapi/src/model.rs – The values the API returns; the layer between entities and GraphQL types.
  • preprocessor/ – Build-time CLI that assembles generated/*.csv from data/*.csv, the GTFS feeds, the Tokyu ODPT JSON, and the MLIT railway data (N02, for track lengths — see Track distances below).
  • data/ – Canonical CSV datasets. Files follow the N!table.csv naming scheme. Detailed instructions are in data/README.md.
  • data_validator/ – CLI that verifies cross-file constraints (cargo run -p data_validator).
  • Makefile – Convenience targets (make help lists them all).

The Worker is the workspace root package. stationapi, preprocessor, and data_validator are native workspace members, so type-checking and linting are split by target (see make check / make clippy).

Tooling and Environment

  • Rust: Use the stable toolchain (rustup default stable) plus the wasm target (rustup target add wasm32-unknown-unknown).
  • worker-build (cargo install worker-build --locked) and wrangler are needed to build and run the Worker.
  • No database. The data is embedded into the WASM binary at build time.
  • Environment variables:
    • ODPT_ACCESS_TOKEN – ODPT consumer key used to download authenticated data such as Seibu Bus GTFS, Keio Bus GTFS, Tokyu Bus JSON, and the Tokyu-operated Ota, Shinagawa, and Meguro community bus GTFS feeds. Only used by preprocessor.
    • DISABLE_BUS_FEATURE – set to true to build rail-only data.
  • Keep local secrets in .env.local (git-ignored) and export them before running make data.
  • make data also downloads the MLIT National Land Numerical Information railway data (N02, about 13 MB, no token) and caches its track GeoJSON under data/N02-25/ (git-ignored). Unlike a bus feed, a failed download fails preprocessor — silently shipping data without track lengths would turn every trackDistanceFromPrevious into null.

Running and Deploying

  • Local development
    1. make data builds generated/*.csv from data/*.csv, the GTFS feeds, the Tokyu ODPT JSON, and N02. Feeds already extracted under data/*-GTFS/ are reused; the ODPT JSON is cached for seven days; N02 is reused while data/N02-25/ exists.
    2. make build compiles the Worker, make dev serves it on http://127.0.0.1:8787.
    3. GET /__ping answers without touching the data, GET /__health reports index sizes, GET / serves GraphiQL, and GET /__schema returns the SDL.
  • Deploying
    • A branch determines the target, and the workflow file encodes it. dev triggers deploy_staging.yml (staging, stationapi-stg); master triggers deploy_production.yml (production, stationapi). No other branch deploys anywhere. build_worker.yml only builds and verifies; it runs on pull requests and on pushes to every branch except dev and master, which the deploy workflows already cover.

    • Never pick the environment with an expression. Unlike if:, a job's environment applies whenever the job runs, so a computed name puts every branch and pull request into that environment's deployment history and hands them every secret it holds, CLOUDFLARE_API_TOKEN included. Each deploy workflow therefore hard-codes one environment, and build_worker.yml declares none. workflow_dispatch has no branch filter, so the deploy jobs also carry an if: pinning them to their branch.

    • The build steps live in .github/actions/build-worker, a composite action all three workflows share, so verification and deployment build identically. A local action needs a checkout first, so each workflow runs actions/checkout with persist-credentials: false (the default leaves GITHUB_TOKEN in .git/config, readable by the third-party code cargo install and npx execute) and then calls the action.

    • A deploy must not ship truncated data. preprocessor only warns when a feed fails, so the deploy workflows pass fail-on-missing-bus-feeds: true and fail on any feed that did not import. Without it, an empty or expired ODPT_ACCESS_TOKEN silently produces a dataset holding only Toei Bus and ships it. build_worker.yml has no environment and therefore no token, so it warns instead — the schema and bundle-size checks do not depend on bus data.

    • Required secrets. CLOUDFLARE_ACCOUNT_ID is a repository secret. CLOUDFLARE_API_TOKEN and ODPT_ACCESS_TOKEN are environment secrets in both staging and production. wrangler cannot mint the API token — it has no such command, and the wrangler login OAuth token lacks the API Tokens Write scope that POST /user/tokens requires — so create it in the dashboard.

    • Scope the API token to what a deploy actually needs. The Edit Cloudflare Workers template is convenient but also grants Workers KV Storage: Edit, Workers R2 Storage: Edit, and Workers Tail: Read, none of which this Worker uses — wrangler.jsonc declares no bindings at all. Build a custom token holding only:

      • Account — Workers Scripts: Edit, Account Settings: Read
      • Zone (trainlcd.app) — Workers Routes: Edit, which registers the custom domains
      • User — User Details: Read, User Memberships: Read

      Do not trim below that. Wrangler resolves the account through the user endpoints, and dropping them surfaces as an opaque code 10000 authentication error rather than a permission message. Setting CLOUDFLARE_ACCOUNT_ID reduces how often wrangler needs the membership lookup but does not remove it.

    • Local deploys use make deploy (staging) and make deploy-production (production). Both refuse to run outside their branch. wrangler is invoked through npx pinned to WRANGLER_VERSION, which appears in the Makefile, in both deploy workflows, and as the composite action's default; keep the four in step. wrangler 4 warns when --env is omitted with multiple environments defined, so the staging target passes --env="" explicitly.

    • wrangler deploy always runs the build.command in wrangler.jsonc (worker-build --release); wrangler offers no flag to skip it. Anything that deploys therefore needs the Rust toolchain, the wasm32 target, and worker-build on PATH.

    • The data lives inside the WASM binary, so a data change needs a rebuild and a redeploy. It is not picked up at runtime.

    • A custom domain cannot be registered twice. When moving a domain, remove it from the old Worker and deploy that first.

Data Management

  • CSV load order depends on the numeric prefix (1!, 2!, ...). When adding datasets, choose a prefix that preserves cross-file dependencies.
  • Column sets live in preprocessor/src/rail.rs (*_COLUMNS). Update them alongside any CSV column change; generated/*.csv must keep the same column order because src/index.rs reads it by name and build.rs by position.
  • Columns whose name starts with # are notes and are not loaded.
  • Through-service junction stations – When a train type runs through a station where its lines connect, add a 5!station_station_types.csv row for every line-specific station_cd at that station, even when those rows share one station_g_cd. The only exception is when the train type explicitly identifies a direction or line-specific operation that excludes one side. Omitting either ID makes the train type selectable from only one line in the app. For example, Hida at Gifu must include both the Takayama Main Line station (1141601) and the Tokaido Main Line station (1150239). Audit both sides whenever adding or editing a through-service pattern.
  • 8!connections.csv holds hand corrections to track lengths. preprocessor computes the length between every pair of adjacent rail stations from N02; a row here (station_cd1, station_cd2, distance in meters, either direction) overrides the computed value. Use it only where N02's geometry or a station's coordinates make the computed value wrong, and cite the source of the corrected value in the pull request.
  • data_validator currently verifies that 5!station_station_types.csv references valid station and type IDs, that 8!connections.csv references valid stations with non-negative distances and no duplicate pair, and that order-sensitive station sequences in 3!stations.csv stay intact under ORDER BY e_sort, station_cd (e.g. the Toei Oedo Line's Tochomae rows, whose misordering silently drops the station from ETA estimation). Extend the validator when new cross-references or order-sensitive spots are introduced and keep the process fail-fast (panic on invalid data).

Testing and Quality

  • Tests – make test runs the unit tests for every native crate, plus cargo test -p stationapi-worker. The Worker only runs on Workers, but src/index.rs is a pure in-memory data structure that builds and executes natively, so its tests (including the grid-versus-full-scan differential check) run here. They need no external services.
  • Type checks – make check covers the native crates and the wasm32 target separately. The Worker also compiles for the host, but only runs on Workers.
  • Linting and formatting – make fmt and make clippy before committing (clippy covers the wasm32 target too). Resolve new Clippy warnings unless an existing #![allow] covers the case.
  • Schema – Changing a GraphQL type changes the SDL. Update schema/public.graphql in the same change; CI compares it against the running Worker's /__schema and fails on any difference. That diff is exactly the client-visible impact.
  • Data verification – Execute cargo run -p data_validator whenever CSVs change and record results in pull requests.
  • IPA coverage audit – Execute make ipa-audit when English or romanized CSV names change. This is a read-only report for data/2!lines.csv, data/3!stations.csv, and data/4!types.csv; it does not fail validation, but highlights unresolved tokens and example names so the IPA dictionary can be extended deliberately.
  • Travel-time benchmark – travel_times/cases.csv lists real travel times (a range and a typical value — the median — in minutes, weekday daytime, trains with the same stopping pattern as the line group) that the arrival estimation is measured against. cargo test -p stationapi-worker (src/travel_times.rs) fails when any case moves more than 1 percentage point further from its typical value than travel_times/baseline.csv records (a range alone hides drifts inside a range widened by one outlier train), when the mean gets worse, or when a recorded case can no longer be estimated. It compares only when the Worker embeds generated/ (the estimation uses track lengths and line groups that exist only there), so build_worker.yml runs it after building the data; on data/*.csv it just prints the table. make travel-time-report (TRAVEL_TIME_API, default the local make dev Worker) measures every case against a Worker built from the generated data. A change to the speed tables, the general speed rules, or the estimation parameters must attach the report from before and after, and must not trade one line's accuracy for the whole set's; after an intended change, run make data and regenerate the baseline with TRAVEL_TIMES_UPDATE_BASELINE=1 cargo test -p stationapi-worker travel_times. Values come from open GTFS feeds (credit them in the README's Data Sources) or from the maintainer.
  • Endpoint benchmarks – make bench (or python3 .claude/skills/benchmark-gql/bench.py) replays every Query field against production (gql.trainlcd.app, script stationapi) and staging (gql-stg.trainlcd.app, script stationapi-stg) and writes a Markdown report under benchmarks/. Both environments embed the same data, so any difference is implementation — which makes this the way to see what a dev-to-master release will do to performance before it ships. Besides client latency it records the Worker's cpuTime, read from wrangler tail --format json and matched to each request by cf-ray; the tail is filtered on a per-run request header, so production's live traffic does not leak into the sample. Collecting CPU time needs the workers_tail (read) scope, and the run sends hundreds of real requests to production — it is not a routine check. Add a case to .claude/skills/benchmark-gql/queries.json whenever a Query field is added, and never edit an existing case's variables: the reports are meant to stay comparable across runs.

GraphQL Query Overview

  • Stations – station, stations, stationGroupStations, stationsNearby, lineStations, stationsByName, lineGroupStations, lineListStations, lineGroupListStations. QueryInteractor enriches stations with lines, companies, station numbers, and train types. lineStations resolves the line's local train-type group (rail kind 0/1 or a priority > 0 type; a bullet-train line, which has no local service, takes the group of its stopping train type with the smallest types.id, e.g. Nozomi or Hayabusa) and returns its stops carrying that group's train type, with or without stationId, so a client that never picked a train type still has the lineGroupId trainRoute requires; when no such group exists — bus lines only carry BusRoute (kind 7, priority 0) variants — it falls back to the line's plain typeless station list so bus stop listings never return empty. stationsByName with fromStationGroupId returns the stations reachable from there: stations sharing a line group with the origin (line_group_cd set, has_train_types true), same-line stations when either side has no line group, and — for rail — stations reachable by transferring, i.e. those for which connectedRoutes with viaLineId set to the station's line returns a route (line_group_cd empty, has_train_types false). The transfer check uses RouteTopology (stationapi/src/domain/route_topology.rs), a time-free copy of the connectedRoutes network built straight from the index without Station entities or time estimates (about 20 ms instead of about 190 ms, cached in its own OnceLock); RouteNetwork holds the same topology, both share trim_pattern and line_group_rows, and a real-data test asserts the two are equal. The check is a ride-limited BFS over line groups (a few ms) plus, only for destinations that are cut vertices of the station–line-group graph, a check that arrives without stopping over at the destination group — otherwise a branch's junction station (Ishibashi-handai-mae on the Minoo Line) would be listed although reaching it on that branch means riding out and back.
  • Lines – line, lines, linesByName. Results include company data and computed line symbols based on repository helpers.
  • Routes – routes, connectedRoutes, estimateArrivalTimes, trainRoute. Paging tokens are currently empty (pagination not implemented).
  • trainRoute – Takes the line group's stops from the repository before any enrichment, slices them to the requested fromStationId–toStationId range (reversing when the request runs backwards), and only then attaches lines, companies, station numbers, train types, and nearby bus routes. Enrichment is per-station and independent, so slicing first does not change any segment; enriching the whole line group first made a three-station request cost the same as a 250-station one. Keep the order — the cost of this query must stay proportional to the requested range, not to the line group. The optional model argument (TrainRouteModel) picks how the segment values are computed. Legacy, the default, keeps the values from when the query was added (#1568): dto::simulation::resolve_speed_profile supplies the speed and acceleration, and arrivalCumulativeMinutes / departureCumulativeMinutes stay null. Older TrainLCD/MobileApp builds run auto mode on these values, so do not change them; Legacy reads frozen copies of the speed tables (domain/legacy_speed_table.rs) so that recalibrating the estimation does not move it. Estimated (TrainLCD/MobileApp's auto mode and GPX generation, so that auto mode, ETA, and GPX take the same time) runs arrival_estimation — the model the speed tables are calibrated against — on the same pre-enrichment slice with the same calibration and distances as estimateArrivalTimes (with legs, it reuses estimate_connected_route_arrival_times, so the values equal estimateArrivalTimes given the same legs). It then replaces each segment's stops, speed, and acceleration with the estimator's and fills the two cumulative fields. A route containing a bus station keeps the Legacy values, because the estimator's kinematics do not apply to buses.
  • Coordinate lookups – index::nearest (k nearest, used by stationsNearby) and index::within_radius (everything inside a radius, used by the nearby-bus-stop enrichment) both go through a per-transport-type grid index (Grid, CSR over 0.05° cells) instead of scanning the whole station table. nearest searches a radius, widens it while fewer than limit stations fall inside, and stops once the radius covers the index — anything outside a radius that already holds limit hits cannot be in the top limit. With transportType omitted it returns rail stations first and bus stops after, each group sorted by distance — the pre-Workers SQL's ORDER BY transport_type, distance. The limit applies to the merged order, so nearest fills it with rail and only asks the bus grid for the remaining slots; a location with limit rail stations returns no bus stops at all. Ties on distance break on station_cd so the order does not depend on an unstable sort. Every station lookup by coordinates runs on every request that enriches rail stations with nearby bus routes, so keep new coordinate queries on the grid rather than adding another full scan.
  • Train types – stationTrainTypes, routeTypes. Train types aggregate by line group and include related lines plus optional train type metadata. Rail variants use TrainTypeKind::{Default, Branch, Rapid, Express, LimitedExpress, HighSpeedRapid, CommuterRapid} (0-6); bus variants use BusRoute (7), which represents a (route_id, shape_id) operation pattern (e.g. 循環 / 短ターン / 支線) generated automatically from the configured GTFS bus feeds (Toei Bus, Seibu Bus, Keio Bus) and the converted Tokyu Bus JSON.
  • Default rail train types – preprocessor fills every active rail line containing at least one station with no station_station_types row with a deterministic, complete all-stop group. The generated rows exist only in generated/*.csv; canonical CSV files remain unchanged. type_cd=100 represents 「普通」 and type_cd=101 represents 「各駅停車」. An existing 100/101 assignment on the line takes precedence; otherwise the label is selected per line through LOCAL_SERVICE_RAIL_LINE_IDS in preprocessor/src/rail.rs. Generated line_group_cd values use 1,000,000,000 + line_cd; generation fails on a collision. Bus lines are excluded and continue to use their GTFS-derived BusRoute groups.
  • GTFS bus integration – preprocessor/src/gtfs/ reads the GTFS feeds into an in-memory representation and then projects them onto the shared stations / lines / types / station_station_types tables (gtfs/integrate.rs). Only routes, stops, trips, and stop_times are read; calendar, shapes, feed_info, and agencies do not affect the output. Every configured GTFS feed is imported, including Seibu Bus and Keio Bus (both downloaded from ODPT with ODPT_ACCESS_TOKEN). Tokyu Bus ordinary-route BusroutePattern, BusstopPole, and BusTimetable JSON are converted into the same representation; pattern IDs become shape_id values so route variants remain queryable as bus TrainTypes. The Tokyu-operated Ota, Shinagawa, and Meguro community buses use their official GTFS feeds and matching JSON routes are excluded to prevent duplicates. ODPT_ACCESS_TOKEN is required for authenticated sources; without it those feeds are skipped with a warning rather than failing the build. Stops whose Tokyu JSON records omit coordinates remain available to name and route queries but not coordinate searches. transport_type (0: rail, 1: bus) on both stations and lines keeps rail and bus records queryable side by side. GTFS IDs are namespaced per feed before import to avoid cross-operator collisions. line_cd (100,000,000+), station_cd / station_g_cd (200,000,000+), and bus type_cd / line_group_cd (100,000,000+) are all deterministic fnv1a hashes that stay clear of the rail data ranges. Disable the entire bus pipeline with DISABLE_BUS_FEATURE=true.
  • Bus stop translations (readings & English) – GTFS-JP translations.txt layouts differ per feed, so load_translations (preprocessor/src/gtfs/parse.rs) resolves columns by header name (Seibu ships 6 columns without record_sub_id; Keio and the Tokyu community feeds ship 7) and indexes each stop_name translation under both keys it may use: record_id (== the stop_id, Seibu — with the "-NN" pole suffix also mapped to the parent stop_id) and field_value (== the Japanese stop_name, Keio / Tokyu community, where record_id is left empty). load_stops then looks a stop's translation up by stop_id first, then by name. Keying only by record_id would silently drop every field_value-keyed feed, leaving station_name_k filled with the kanji stop_name and station_name_r empty. Readings arriving as half-width katakana (ニシハチオウジ, Keio / Tokyu community) are folded to full-width via romaji::to_fullwidth_katakana() before storage.
  • Bus English-name fallback – When a feed provides no English (en) translation for a stop — e.g. Tokyu Bus ordinary-route JSON, which carries only dc:title and odpt:kana — stationapi/src/domain/romaji.rs::romaji_display_name() derives a modified-Hepburn romanization (with macrons for long vowels, matching the curated rail style: Tōkyō / Kyōto / Shin-Ōsaka) from the kana reading, and the GTFS reader fills stop_name_r with it. The fallback never overwrites a real en value, and a reading with no convertible kana stays NULL rather than emitting a partial transcription. Because stop_name_r is the single upstream source that fans out into the stations projection, search_by_name, and the romanized bus route/headsign names, this supplements every English-facing surface at once. When projecting into stations, station_name_rn is filled with the plain-ASCII spelling via romaji::strip_macrons() (Tōkyō → Tokyo), mirroring the rail dataset's _r (macron) / _rn (macron-free) column pair.
  • Track distances – Station.trackDistanceFromPrevious is the track length in meters from the station before it in the returned list, for clients that total a ride's distance (TrainLCD/MobileApp's ride log). It is distinct from trainRoute's distanceFromPrevious, which stays the straight-line distance the app's running simulation relies on. The arrival estimation (estimateArrivalTimes and trainRoute with Estimated) uses the track length as the running distance (straight line × detour factor where none exists) together with the recalibrated speed tables (speed_table / segment_speed_table, SpeedCalibration::Recalibrated); scripts/compute_speed_table.py fits those tables with the same distances, so it needs generated/ from make data. The ride times of connectedRoutes and trainRoute with Legacy keep the original estimation (straight line × detour factor and the frozen tables in domain/legacy_speed_table.rs, the EstimationParams default), so recalibrating never moves route search results. Route search ride times therefore differ from the ETA. Only lineStations, lineGroupStations, trainRoute, and stations(ids) fill it (attach_track_distances in QueryInteractor, called once the order is final). stations(ids) returns the stations in the order of ids — the repository keeps that order and enrichment never reorders — so a client passing a route's station IDs (MobileApp's sids deep links) gets the lengths between IDs adjacent in that order; a pair that is not adjacent in the track data (an ID list skipping stations) is null. The first station, the first station of each trainRoute leg, sections without N02 geometry, and every other query return null, and clients fall back to the straight line there. preprocessor/src/track/ builds an undirected graph from N02's RailroadSection LineStrings, collects every pair adjacent in a line's (e_sort, station_cd) order or a line group's station_station_types.id order (with and without closed stations, plus the seam of loop services within 3 km), snaps each station to every track within 500 m, and takes the Dijkstra path minimizing 2 × snap offset + track length — snapping to the nearest track alone picks another line of the same operator at large stations (Honmachi) and detours through a transfer station. It first measures on the operators of both stations' companies (OPERATOR_ALIASES maps companies whose name differs from N02's), then retries on every operator for lines running on another company's track; a path longer than max(3 × straight, straight + 5 km) counts as unmeasured, and a result shorter than the straight line is raised to it. Pairs in one station group are 0. build.rs writes the sorted pairs to connections.bin, and index::track_distance binary-searches it, so no index is built at isolate start. When bumping the N02 edition, change the URL and cache directory in preprocessor/src/track/mod.rs and the Cache N02 railway data step in .github/actions/build-worker/action.yml together and compare the unmeasured count in the preprocessor log. docs/architecture.md (駅間の線路の長さ) has the details.
  • TTS metadata – Station, StationNested, Line, LineNested, TrainType, and TrainTypeNested expose name_ipa / name_roman_ipa plus name_tts_segments for multi-segment pronunciation output. Use name_tts_segments when clients need per-token SSML construction for mixed-language names such as Kasai-Rinkai Park.
  • Connected routes – connectedRoutes finds transfer routes automatically, like a journey planner, using a frequency-based RAPTOR search in stationapi/src/domain/route_search.rs. Each rail line group is a pattern, station groups are the transfer nodes, and ride times come from arrival_estimation; bus lines are excluded. The cost adds a per-boarding wait by TrainTypeKind (limited express 15 min, express / high-speed rapid 5 min, others 3 min) and a 3-minute transfer walk — without the wait, infrequent limited expresses would beat the Yamanote Line. Rounds give the time/transfer Pareto set; alternatives come from re-searching with one leg's parallel line groups banned along that leg (at most 8 searches), and are dropped beyond 1.15 × best + 15 min or with two more transfers than the Pareto set. Alternative routes that stop at the same station group in two different legs (backtracking to re-board a banned train) are dropped; pass-through stations are not counted, and the Pareto routes of the first search are never dropped this way (otherwise a station stationsByName reports as reachable could get no route). Results are ranked by cost + 5 min per transfer and capped at 6. The time and transfer count stay internal (the API does not return them), so ordering happens on the server: sortBy: ConnectedRouteSort picks Recommended (the ranking above, the default when omitted), ArrivalTime (the search's own ride time — excluding the first train's wait, and estimated with the original calibration, so it differs from what estimateArrivalTimes reports — then fewer transfers), or TransferCount (fewer transfers, then earlier arrival). route_search::sort_journeys only reorders the set search returned, with a stable sort so ties keep the recommended order; the set itself never depends on sortBy. The network (every rail line group plus its time estimates, about 190 ms natively) is built lazily into a OnceLock by StationRepository::get_route_network on the first connectedRoutes call, so other queries never pay for it. Each route is a list of legs shaped for the app's one-train-at-a-time flow: every leg carries its boarding and alighting Station, both on the line of the line group the search rode — so at a transfer the previous leg's alighting station and the next leg's boarding station may be different stations of one station group — stationGroupIds, the station groups from boarding to alighting in travel order including pass-through stations (the search's JourneyLeg.station_group_ids as-is — station groups rather than station IDs so the client can match them against whichever train type it picks, which may run on another line; a group appears twice on patterns such as the Oedo Line's Tochomae), and trainTypes, every train type usable on that leg (real groupIds, so the client picks one and calls lineGroupStations): it is exactly routeTypes(boarding station group, alighting station group, alighting station's line) (same dedup, same lines, same order — the use case calls get_train_types), because the search collapses parallel services such as local and rapid into one route and the app needs them to list types and default to the local. viaLineId, like routeTypes, is the line of the tapped search result and keeps only routes whose last leg arrives on that line. estimateArrivalTimes and trainRoute accept legs: [RouteLegInput!] (the groupId of the train type picked from each leg's trainTypes, plus the leg's fromStation.id and toStation.id) and then return values for the whole transfer route. A leg endpoint missing from the chosen line group is matched by station group (the picked local may stop at another line's station of the same group), preferring an exact station_cd and, among same-group candidates, the pair giving the shortest slice (through services list two stations of a junction group). ETA estimates each leg on its own line group only and chains them from the origin, adding the 3-minute walk and the next train type's wait at each transfer (the same allowance the search ranks by), returning one route with an empty id; trainRoute concatenates each leg's segments (each leg restarts at distance 0). Both slice legs with the same function, taking the shorter arc on loop lines, so their station sequences match. More than MAX_RIDES (6) legs — more than connectedRoutes ever returns — legs that do not connect, ends that differ from fromStationId / toStationId, or combining legs with viaLineIds / directionId / lineGroupId are errors. docs/architecture.md (乗換経路探索) has the details, and docs/route-search.md documents the search internals (data structures, the scan, pruning, alternatives, determinism).
  • Changes to the published contract require coordinated updates to schema/public.graphql, the async-graphql types in src/graphql/, and, when the shape of a value changes, stationapi/src/model.rs and the DTO conversions.

Version Control (Git)

Version control is plain Git. gh is the tool for pull requests, and GitHub Actions consumes the same refs.

  • origin/dev is the base for ordinary work. Fetch it before branching so a feature branch does not start from a stale dev; master is release-only.
  • Stage deliberately. git add -u picks up edits to tracked files; add an untracked file by explicit path rather than git add -A / git add ., so scratch files do not ride along. Read git status before committing.
  • Never rewrite a pushed commit without asking. git commit --amend, git rebase, and anything that needs git push --force-with-lease rewrite published history. Confirm with the user first.
  • Update a remote-tracking ref with an explicit refspec. git fetch origin dev leaves refs/remotes/origin/dev to whatever remote.origin.fetch happens to be; a narrowed refspec on one machine silently leaves origin/dev stale, and every branch, rebase, and diff taken from it is then measured against an old commit.
  • The stash stack is shared with every worktree of this repository. Prefer a temporary WIP commit to set work aside; if you must stash, use git stash push -u -m "<tag>" and restore with git stash apply <sha> rather than a bare git stash pop.

A typical change:

git fetch origin "+refs/heads/dev:refs/remotes/origin/dev"  # explicit refspec, so origin/dev cannot be stale
git switch -c feature/<description> origin/dev
# ... edit files ...
git status  # confirm exactly what the change contains
git add -u  # plus explicit paths for new files
git commit -m "日本語の単文"
git push -u origin feature/<description>

Commands this guide and .claude/skills/create-pr rely on:

Purpose Command
Repository root git rev-parse --show-toplevel
Working-tree state git status --short
Current branch git rev-parse --abbrev-ref HEAD
Commit subjects on a branch git log --pretty=%s origin/dev..origin/<branch>
Files changed against a base git diff --name-only origin/dev..origin/<branch>
Rebase onto the latest dev git fetch origin "+refs/heads/dev:refs/remotes/origin/dev" && git rebase origin/dev

CONTRIBUTING.md documents the same workflow for outside contributors. Keep the two aligned — base branch, naming convention, and pull-request rules are identical.

Contribution Guidelines

  • Git-flow – Follow Git-flow with dev serving as this repository's develop branch. Create ordinary work branches from the latest origin/dev, use the feature/<description> naming convention, and target their pull requests to dev. Do not create or target a branch named develop. Version Control (Git) above has the full command sequence.
  • Pull requests – Assign every pull request to @TinyKitten when creating it, open it as ready for review rather than as a draft, and use .github/pull_request_template.md without omitting or replacing its sections or checklists.
  • Prioritize quality and performance over implementation speed – Always favor code quality and runtime performance over velocity. Be mindful of algorithmic complexity and look for opportunities to replace O(n×m) linear scans with O(n+m) indexed lookups (e.g., HashMaps). The indexes are rebuilt on every isolate start and every request scans them, so prefer indexed lookups over repeated full scans. When a change affects performance, document the before/after complexity and query plan impact in the pull request.
  • Document the commands you executed (for example, make fmt && make clippy && make test) and their outcomes in every pull request.
  • For data pipeline or schema updates, add architectural notes under docs/ and synchronize README references so onboarding materials stay accurate.
  • When modifying QueryInteractor, ensure the enrichment steps (companies, train types, line symbols) still behave as expected. Double-check helper methods such as update_station_vec_with_attributes and build_route_tree_map.
  • Introducing new tables, endpoints, or feature flags must come with matching updates to this document and any other affected guidance.

Maintenance

Keep this guide aligned with the repository. If a workflow, environment requirement, or endpoint changes, update AGENTS.md in the same pull request so automation agents and contributors work from current instructions.