A lightweight HTTP API for fast substring lookup against one or more prebuilt search index files.
The service loads JSON-based search data into memory at startup and exposes a simple query endpoint for autocomplete-style lookup and search integrations.
- API docs UI: https://lookup.artsdatabanken.no/
- Runtime: Node.js
- Framework: Express
- Search engine:
indexed-substring-search
- Fast in-memory substring search
- Simple HTTP API with JSON responses
- Swagger UI available at
/and/api-docs - Supports loading multiple index files from the
data/directory - Dockerized runtime for deployment
src/index.js- application entrypoint and server bootstrapsrc/routes.js- API route definitionssrc/lookupIndex.js- index loading and query logicsrc/swagger.js/src/swagger.json- Swagger UI setup and API definitiondata/- search index files and fallback responsesDockerfile- production container imageazure-pipelines-*.yml- CI/CD pipelines for Docker image build and push
- Node.js
- npm
- Optional: Docker
Recommended: use a Node.js version compatible with the Docker image. The repository currently builds with Node 21 in Docker.
Install dependencies from the project root:
npm installStart the API directly with Node:
node src/index.js --port 9876 --dataPath ./data/Then open:
- Swagger UI:
http://localhost:9876/ - Swagger UI alt path:
http://localhost:9876/api-docs - Query endpoint example:
http://localhost:9876/v1/query?q=lecanorales
Run the app with nodemon for automatic restart on JavaScript file changes:
npm run devThis starts:
node src/index.js --port 9876 --dataPath ./data/GET /v1/query?q=<text>
Example:
curl "http://localhost:9876/v1/query?q=lecanorales"Example response shape:
{
"query": "lecanorales",
"result": [
{
"score": 9.13,
"title": "Kantlavordenen",
"url": "Biota/Fungi/Ascomycota/Pezizomycotina/Lecanoromycetes/Lecanorales"
}
]
}- Missing or empty queries return the contents of
data/emptyquery.json - Queries with no matches return the contents of
data/nohits.json - The API returns up to 20 results
The index is built from files in data/ whose names contain full-text-index.
The current implementation supports:
.jsonfiles containing an object with anitemsfield- non-
.jsonfiles treated as newline-delimited JSON entries
Each entry contains:
Object containing the payload returned by the API when one or more search terms match.
Map where:
- key = score for a hit on the listed terms, where higher is better
- value = array of search terms associated with that score
Example:
{
"hit": {
"title": "Kantlavordenen",
"url": "Biota/Fungi/Ascomycota/Pezizomycotina/Lecanoromycetes/Lecanorales"
},
"text": {
"587": [
"AR-128057",
"Pezizomycotina",
"Ekte sekksporesopper",
"Ekte sekksporesoppar"
],
"652": ["AR-127735", "Lecanoromycetes", "Kantlaver", "Kantlavar"],
"913": ["Lecanorales", "Kantlavordenen", "Kantlavordenen"],
"932": ["AR-1001"]
}
}Expected supporting files in data/:
emptyquery.json- fallback response whenqis missing or emptynohits.json- fallback response when no search results are foundfull-text-index*- one or more search index source files
Build the image:
docker build -t generic-substring-lookup-api .Run the container:
docker run --rm -p 9876:9876 -v "$(pwd)/data:/data" generic-substring-lookup-apiThe container starts the service with:
node --max_old_space_size=8192 src/index.js --port 9876 --dataPath /data/Current test command:
npm testAt the moment this is only a placeholder and does not run an automated test suite.
The repository includes Azure Pipelines configuration:
azure-pipelines-test.yml- builds and pushes the Docker image taggedteston thetestbranchazure-pipelines-prod.yml- builds and pushes the Docker image taggedlateston themasterbranch
There is also a deploy.sh helper that logs in to Docker Hub, pushes the image, and posts a Slack notification for master.
A few implementation details are worth checking before making operational assumptions:
- The no-hit check in
src/lookupIndex.jsusesif (!Object.keys(result)), which may not correctly detect an empty result depending on the returned data type.
Check that:
data/emptyquery.jsonexists and is valid JSONdata/nohits.jsonexists and is valid JSON- at least one file containing
full-text-indexexists indata/ - the search input structure matches the expected schema
- verify that the query parameter is named
q - test with exact known terms from the source data
- inspect the loaded
full-text-index*files
The index is built fully in memory. For large datasets:
- expect significant startup memory usage
- consider increasing runtime memory limits
- reduce or split dataset size if needed
When updating this service:
- Keep route handlers minimal and put search logic in
src/lookupIndex.js - Preserve the current CommonJS module style unless intentionally refactoring the whole codebase
- Update Swagger documentation if API behavior changes
- Validate changes with representative data files
- Verify Docker behavior if changing startup or runtime assumptions
- Express: https://expressjs.com/
- Swagger UI Express: https://www.npmjs.com/package/swagger-ui-express
- indexed-substring-search: https://www.npmjs.com/package/indexed-substring-search
- Azure Pipelines Docker task: https://learn.microsoft.com/azure/devops/pipelines/tasks/reference/docker-v2