Skip to content

About

AcoustID's web site and API

Topics

Resources

Stars

91 stars

Watchers

6 watching

Forks

Latest commit

 

History

1,379 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AcoustID Server

This software is only meant to run on acoustid.org. Running it on your own server is not supported. It's possible, but you need to understand the system well enough and even then it's probably not going to be useful to you.

Local Development

You need Python 3.12 or newer to run the code. On Ubuntu, you can install the required packages using the following command:

sudo apt install python3 python3-dev python3-venv

You also need uv, see their installation instructions.

Setup Python virtual environment:

uv sync
source .venv/bin/activate

Start the required services using Docker:

export COMPOSE_FILE=docker-compose.yml:docker-compose.localhost.yml

docker compose up -d redis postgres index

Prepare the configuration file:

cp acoustid.conf.dist acoustid.conf
vim acoustid.conf

Initialize the local database:

uv run alembic upgrade head

Run the applications:

python manage.py run web
python manage.py run api
python manage.py run cron
python manage.py run worker

Local Testing

You can use the provided docker-compose.yml file to quickly set up a test environment:

export COMPOSE_DOCKER_CLI_BUILD=1
export DOCKER_BUILDKIT=1
export COMPOSE_FILE=docker-compose.yml:docker-compose.localhost.yml
export COMPOSE_PROFILES=frontend,backend

docker-compose up -d

Public data export

Writes the daily incremental files published at data.acoustid.org into a local directory:

python manage.py data export --directory /var/lib/acoustid/data-export

The layout is {YYYY}/{YYYY-MM}/{YYYY-MM-DD}-{table}.jsonl.gz, gzipped JSON Lines with null fields removed, seven files per day. Only complete days are exported, so the newest file is always for yesterday.

Every directory also carries an index.html and an index.json listing its contents, which is the only way the archive can be browsed at all: whatever serves it may have no directory listing of its own, and a bucket certainly does not. Both are part of the published interface down to the bytes, so they are reproduced exactly -- 1024-based sizes with one decimal in the HTML, compact JSON with a size per file and a trailing slash per directory. nginx's autoindex is not a substitute: it answers at the directory URI rather than at index.json, names its fields differently, drops the trailing slash and reports an mtime that changes every time a file is rewritten.

The export brings those up to date for every directory it goes through, at the end of the run. Unlike a data file, an index that is already there is rewritten, because its directory's contents are the one thing that can have changed -- but only when the bytes actually differ, so an unchanged directory is left alone entirely.

A day is exported an hour after it ends at the earliest, and only once no transaction that started before the day ended is still running. created and updated are transaction start time but a row only becomes visible at commit, so a transaction straddling midnight would otherwise leave the file short by exactly its rows -- permanently, since a file that exists is never regenerated. A day that is not ready yet is skipped without writing anything and picked up by a later run.

Reading that requires the export user to be able to see other sessions:

GRANT pg_read_all_stats TO acoustid;

Without it pg_stat_activity reports NULL for other sessions and every day would look settled, so the command refuses to start rather than exporting on the strength of a check that cannot fail.

Runs walk backwards from midnight today and skip files that already exist, so running this hourly is safe and re-running it with a larger window is how a gap gets backfilled:

python manage.py data export --max-days 45

The default window is 30 days. The directory can also come from [export] directory in acoustid.conf or from ACOUSTID_EXPORT_DIRECTORY.

Getting the files from that directory to the bucket that serves data.acoustid.org is a separate step. That sync must be additive. The export directory only ever holds the last --max-days of files, while the bucket holds every file back to 2011, so an rsync --delete out of it would delete the archive.

An index describes the directory it is written in, which makes that same gap a second hazard: an index generated from a directory holding only the last 30 days lists only those files, and syncing it over a complete month in the bucket replaces a full listing with a partial one -- no files lost, but they stop being findable. So either the local directory is the whole archive, or the index files are kept out of the sync and built where the whole archive is. Serving the directory directly, rather than syncing it anywhere, makes the question moot.

Rebuilding every index in a tree, without exporting anything:

python manage.py data generate-indexes --directory /var/lib/acoustid/data-export

That walks what is already on disk and reads the sizes from the files themselves, so it needs no database and does not care which process wrote the tree. It is how an archive exported before the index files existed gets its listings, and it is safe to re-run: an index whose bytes come out the same is not rewritten.

Database migrations

Upgrading the database schema online:

alembic upgrade head

Upgrading the database schema offline:

alembic upgrade <previous-rev>:head --sql

Generating a new database schema change:

alembic revision --autogenerate -m "my message"

Unit tests

Before you can run the test suite, you need to create a configuration file called acoustid-test.conf. This should use a separate database from the one you use for development, but it should have the same structure.

You can then run the test suite like this:

pytest -v tests/

The first thing it does is setting up the database. Normally you shouldn't need to do this more than once, so the next time you can run the test suite without the database setup code:

SKIP_DB_SETUP=1 pytest -v tests/

About

AcoustID's web site and API

Topics

Resources

Stars

91 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages