Skip to content

Skip the transliteration backend for ASCII text - #201

Merged
un33k merged 1 commit into
un33k:masterfrom
rafaelborja:ascii-fast-path
Sep 27, 2026
Merged

un33k merged 1 commit into
un33k:masterfrom
rafaelborja:ascii-fast-path

Conversation

@rafaelborja

Copy link
Copy Markdown
Contributor

Why

Every slug goes through _transliterate(), and the first call imports the transliteration backend even when the input is plain ASCII. With the default backend, text-unidecode, that import decodes its whole replacement table: about 3 MB of strings that then stay in memory for the life of the process.

For plain ASCII input the table is never needed. All three supported backends (text-unidecode, Unidecode, anyascii) map every 7-bit ASCII codepoint to itself; I checked all 128 on each. Most real-world slugs are plain ASCII, so most processes pay this memory for nothing.

Change

Two lines at the top of _transliterate(), in both slugify.py and _legacy.py:

if text.isascii():
    return text

ASCII input is returned unchanged without importing a backend. Non-ASCII input takes exactly the same path as before.

Memory impact

CPython 3.14, fresh process, text-unidecode backend:

master this PR
import slugify 0.84 MiB traced 0.86 MiB
+ slugify("Living Room Light") 4.16 MiB traced / +11.0 MiB RSS 0.86 MiB / +2.5 MiB RSS
+ a non-ASCII slug 4.16 MiB 4.17 MiB (unchanged)

On a real application: Home Assistant slugifies entity, device and area names at startup, and nearly all of them are ASCII. Measured on an arm64 HAOS install, with this change applied to the pinned release, Home Assistant's resident memory dropped by about 5 MB (anonymous memory by about 6 MB) for the whole time it runs. That was with lazy imports already enabled, and text-unidecode was no longer imported at all. On small devices such as a Raspberry Pi with 1–2 GB of RAM, that's permanent headroom at no cost.

Risks and how they are covered

  • Output could change for some input. The frozen 2,688-case legacy differential in test_release.py still passes. A separate randomized differential of 224,448 cases gives byte-identical output to master: both algorithms × auto/text-unidecode/unidecode/anyascii × 7 option sets × 4,008 inputs, including HTML entities, quotes, numbers, CJK, Cyrillic and control characters. On the Home Assistant install above, all 2,988 real entity, device and area names, plus every 7-bit ASCII character, gave identical output.
  • HTML entities such as é become non-ASCII once decoded. They're decoded before _transliterate() runs, so they still reach the backend. The randomized differential covers them.
  • Invalid backend names are still rejected. The backend value is validated before _transliterate() in both implementations, so the early return can't hide a bad value.
  • _legacy.py says not to modify it. This change doesn't alter its output: the frozen legacy differential exists to guard exactly that, and it passes. If you'd rather keep the file untouched, I can limit the change to slugify.py.
  • One intentional visible difference: if a backend is installed but broken (for example it raises ModuleNotFoundError for one of its own dependencies), or no backend is installed at all, ASCII input now slugifies instead of raising. Three assertions in test_backend_selection_and_no_fallback_on_broken_install used 'x' to exercise backend selection. They now use 'é', so they still test the same thing. A new test checks that ASCII input doesn't import a backend under either algorithm.

Tests

126 passed, with and without the optional backends installed.

Every backend (text-unidecode, Unidecode, anyascii) maps 7-bit ASCII to
itself, so ASCII input no longer imports one. With text-unidecode this
avoids loading its ~3 MB table for the common case of ASCII slugs.
Output is unchanged, including the frozen legacy differential.
@un33k

un33k commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Thanks Rafael — verified and accepting as-is.

What I checked locally against your branch:

  • Full suite: 126 passed
  • Frozen 2,688-case legacy differential passes
  • Independent 560-case corpus (both algorithms x all backends x entities/unicode/limits/separators) byte-identical to master
  • Exhaustive all-128 ASCII check byte-identical
  • Fresh-process check: ASCII slugs no longer import a backend under either algorithm; non-ASCII imports exactly as before
  • Invalid backend names still rejected for ASCII input
  • mypy, repo-configured pycodestyle/flake8, git diff --check clean

Decision: taking the two-line fast path in both files. The legacy output is provably unchanged, which is exactly what the frozen-output policy's non-behavioral exception is for — and limiting this to modern-only would withhold the saving from the default path where it matters most (your Home Assistant measurement makes that concrete). The intentional broken/missing-backend edge (ASCII now slugifies instead of raising) reads as a fix, and your updated tests cover it.

Squash-merging with your authorship preserved, then shipping as 9.1.2.

🚀 Generated with Dojo ⛩️

@un33k
un33k merged commit cf6cd3d into un33k:master Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants