This software is only meant to run on acoustid.org. Running it on your own server is not supported. It's possible, but you need to understand the system well enough and even then it's probably not going to be useful to you.
You need Python 3.12 or newer to run the code. On Ubuntu, you can install the required packages using the following command:
sudo apt install python3 python3-dev python3-venv
You also need uv, see their installation instructions.
Setup Python virtual environment:
uv sync
source .venv/bin/activate
Start the required services using Docker:
export COMPOSE_FILE=docker-compose.yml:docker-compose.localhost.yml
docker compose up -d redis postgres index
Prepare the configuration file:
cp acoustid.conf.dist acoustid.conf
vim acoustid.conf
Initialize the local database:
uv run alembic upgrade head
Run the applications:
python manage.py run web
python manage.py run api
python manage.py run cron
python manage.py run worker
You can use the provided docker-compose.yml file to quickly set up a test environment:
export COMPOSE_DOCKER_CLI_BUILD=1
export DOCKER_BUILDKIT=1
export COMPOSE_FILE=docker-compose.yml:docker-compose.localhost.yml
export COMPOSE_PROFILES=frontend,backend
docker-compose up -d
Writes the daily incremental files published at data.acoustid.org into a local directory:
python manage.py data export --directory /var/lib/acoustid/data-export
The layout is {YYYY}/{YYYY-MM}/{YYYY-MM-DD}-{table}.jsonl.gz, gzipped JSON
Lines with null fields removed, seven files per day. Only complete days are
exported, so the newest file is always for yesterday.
Every directory also carries an index.html and an index.json listing its
contents, which is the only way the archive can be browsed at all: whatever
serves it may have no directory listing of its own, and a bucket certainly does
not. Both are part of the published interface down to the bytes, so they are
reproduced exactly -- 1024-based sizes with one decimal in the HTML, compact
JSON with a size per file and a trailing slash per directory. nginx's
autoindex is not a substitute: it answers at the directory URI rather than at
index.json, names its fields differently, drops the trailing slash and
reports an mtime that changes every time a file is rewritten.
The export brings those up to date for every directory it goes through, at the end of the run. Unlike a data file, an index that is already there is rewritten, because its directory's contents are the one thing that can have changed -- but only when the bytes actually differ, so an unchanged directory is left alone entirely.
A day is exported an hour after it ends at the earliest, and only once no
transaction that started before the day ended is still running. created and
updated are transaction start time but a row only becomes visible at commit,
so a transaction straddling midnight would otherwise leave the file short by
exactly its rows -- permanently, since a file that exists is never regenerated.
A day that is not ready yet is skipped without writing anything and picked up
by a later run.
Reading that requires the export user to be able to see other sessions:
GRANT pg_read_all_stats TO acoustid;
Without it pg_stat_activity reports NULL for other sessions and every day
would look settled, so the command refuses to start rather than exporting on
the strength of a check that cannot fail.
Runs walk backwards from midnight today and skip files that already exist, so running this hourly is safe and re-running it with a larger window is how a gap gets backfilled:
python manage.py data export --max-days 45
The default window is 30 days. The directory can also come from [export] directory in acoustid.conf or from ACOUSTID_EXPORT_DIRECTORY.
Getting the files from that directory to the bucket that serves
data.acoustid.org is a separate step. That sync must be additive. The
export directory only ever holds the last --max-days of files, while the
bucket holds every file back to 2011, so an rsync --delete out of it would
delete the archive.
An index describes the directory it is written in, which makes that same gap a second hazard: an index generated from a directory holding only the last 30 days lists only those files, and syncing it over a complete month in the bucket replaces a full listing with a partial one -- no files lost, but they stop being findable. So either the local directory is the whole archive, or the index files are kept out of the sync and built where the whole archive is. Serving the directory directly, rather than syncing it anywhere, makes the question moot.
Rebuilding every index in a tree, without exporting anything:
python manage.py data generate-indexes --directory /var/lib/acoustid/data-export
That walks what is already on disk and reads the sizes from the files themselves, so it needs no database and does not care which process wrote the tree. It is how an archive exported before the index files existed gets its listings, and it is safe to re-run: an index whose bytes come out the same is not rewritten.
Upgrading the database schema online:
alembic upgrade head
Upgrading the database schema offline:
alembic upgrade <previous-rev>:head --sql
Generating a new database schema change:
alembic revision --autogenerate -m "my message"
Before you can run the test suite, you need to create a configuration file called acoustid-test.conf. This should use a separate database from the one you use for development, but it should have the same structure.
You can then run the test suite like this:
pytest -v tests/
The first thing it does is setting up the database. Normally you shouldn't need to do this more than once, so the next time you can run the test suite without the database setup code:
SKIP_DB_SETUP=1 pytest -v tests/