Skip the transliteration backend for ASCII text - #201
Merged
Merged
Conversation
Every backend (text-unidecode, Unidecode, anyascii) maps 7-bit ASCII to itself, so ASCII input no longer imports one. With text-unidecode this avoids loading its ~3 MB table for the common case of ASCII slugs. Output is unchanged, including the frozen legacy differential.
Owner
|
Thanks Rafael — verified and accepting as-is. What I checked locally against your branch:
Decision: taking the two-line fast path in both files. The legacy output is provably unchanged, which is exactly what the frozen-output policy's non-behavioral exception is for — and limiting this to modern-only would withhold the saving from the default path where it matters most (your Home Assistant measurement makes that concrete). The intentional broken/missing-backend edge (ASCII now slugifies instead of raising) reads as a fix, and your updated tests cover it. Squash-merging with your authorship preserved, then shipping as 9.1.2. 🚀 Generated with Dojo ⛩️ |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Every slug goes through
_transliterate(), and the first call imports the transliteration backend even when the input is plain ASCII. With the default backend, text-unidecode, that import decodes its whole replacement table: about 3 MB of strings that then stay in memory for the life of the process.For plain ASCII input the table is never needed. All three supported backends (text-unidecode, Unidecode, anyascii) map every 7-bit ASCII codepoint to itself; I checked all 128 on each. Most real-world slugs are plain ASCII, so most processes pay this memory for nothing.
Change
Two lines at the top of
_transliterate(), in bothslugify.pyand_legacy.py:ASCII input is returned unchanged without importing a backend. Non-ASCII input takes exactly the same path as before.
Memory impact
CPython 3.14, fresh process, text-unidecode backend:
import slugifyslugify("Living Room Light")On a real application: Home Assistant slugifies entity, device and area names at startup, and nearly all of them are ASCII. Measured on an arm64 HAOS install, with this change applied to the pinned release, Home Assistant's resident memory dropped by about 5 MB (anonymous memory by about 6 MB) for the whole time it runs. That was with lazy imports already enabled, and text-unidecode was no longer imported at all. On small devices such as a Raspberry Pi with 1–2 GB of RAM, that's permanent headroom at no cost.
Risks and how they are covered
test_release.pystill passes. A separate randomized differential of 224,448 cases gives byte-identical output to master: both algorithms ×auto/text-unidecode/unidecode/anyascii× 7 option sets × 4,008 inputs, including HTML entities, quotes, numbers, CJK, Cyrillic and control characters. On the Home Assistant install above, all 2,988 real entity, device and area names, plus every 7-bit ASCII character, gave identical output.ébecome non-ASCII once decoded. They're decoded before_transliterate()runs, so they still reach the backend. The randomized differential covers them.backendvalue is validated before_transliterate()in both implementations, so the early return can't hide a bad value._legacy.pysays not to modify it. This change doesn't alter its output: the frozen legacy differential exists to guard exactly that, and it passes. If you'd rather keep the file untouched, I can limit the change toslugify.py.ModuleNotFoundErrorfor one of its own dependencies), or no backend is installed at all, ASCII input now slugifies instead of raising. Three assertions intest_backend_selection_and_no_fallback_on_broken_installused'x'to exercise backend selection. They now use'é', so they still test the same thing. A new test checks that ASCII input doesn't import a backend under either algorithm.Tests
126 passed, with and without the optional backends installed.