Skip to content

Support uppercase hexadecimal references in modern slugs - #195

Merged
un33k merged 1 commit into
un33k:masterfrom
rupayon123:contribution/modern-uppercase-hex-20260913
Sep 18, 2026
Merged

un33k merged 1 commit into
un33k:masterfrom
rupayon123:contribution/modern-uppercase-hex-20260913

Conversation

@rupayon123

Copy link
Copy Markdown
Contributor

Change

Add uppercase-X hexadecimal HTML reference support to the explicitly selected modern algorithm. For example, slugify('A', algorithm='modern') now returns a rather than x41. HTML accepts both x and X in this position: https://html.spec.whatwg.org/multipage/syntax.html#character-references

The default/explicit legacy algorithm retains its existing output and lowercase-x pattern, following DOJO.md's compatibility boundary. The hexadecimal opt-out still works. Invalid references continue to be handled independently, so a malformed reference does not prevent a valid neighbor from decoding.

Validation

Three new regression failures reproduced the missing modern support before the change. Afterward:

  • Python 3.11.15 with text-unidecode 1.3: 109 tests and 32 subtests pass, including the existing 2,688-case legacy differential check and real CLI execution.
  • mypy: no issues in five source files.
  • Repository pycodestyle and flake8 commands pass; git diff --check passes.
  • Original test.py remains unchanged; new tests are in test_release.py.

The full interpreter/backend tox matrix and packaging/release checks were not run locally. No release or universal output-compatibility claim is made. Prepared with AI assistance; tests were executed locally.

@un33k

un33k commented Sep 18, 2026

Copy link
Copy Markdown
Owner

Confirmed and merging. This adds uppercase-X hexadecimal reference support to the explicitly opted-in modern algorithm only; the default/legacy path keeps its lowercase-x pattern and existing x41 output, and the hexadecimal opt-out and independent invalid-reference handling are preserved. HTML does accept both x and X in that position, so this is a correct, compatibility-safe improvement. Thanks, @rupayon123.

🚀 Generated with Dojo ⛩️

@un33k
un33k merged commit df37f92 into un33k:master Sep 18, 2026
un33k added a commit that referenced this pull request Sep 18, 2026
Bump to 9.1.0. Modern-only: decode uppercase &#X..; hex references
(Rupayon Haldar, #195); preserve fitting post-replacement output during
truncation (emme1t, #193); relocate up-front argument type validation to
the modern path (Jon Bailey, #196). Fix add_uppercase_char atomicity
(Cristian Ramirez, #194). Legacy output unchanged.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️
un33k added a commit that referenced this pull request Sep 18, 2026
* Split frozen legacy pipeline into slugify/_legacy.py

Move the legacy slug pipeline into a dedicated frozen module and make the
public slugify() a thin dispatcher: algorithm='legacy' (default) calls the
frozen _legacy implementation, algorithm='modern' runs the modern pipeline.
Relocate the existing up-front TypeError validation for bool/non-int
max_length and non-str separator (from #196) into the modern path only, so
legacy output is unchanged while modern keeps rejecting invalid types.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Reorganize tests under tests/ with frozen legacy suite

Move the test suite into tests/: the original upstream legacy suite becomes
the frozen tests/test_legacy.py (contents unchanged), mirroring the _legacy.py
code split, alongside tests/test_release.py and the add_uppercase test. The
modern-only bool/separator validation test moves to the modern suite. Update
pyproject testpaths, MANIFEST.in, and tox commands to the tests/ layout.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Document legacy-frozen policy and split in DOJO.md and README

Record that legacy is architecturally frozen (slugify/_legacy.py and
tests/test_legacy.py) and that all new work targets algorithm='modern'.
Add a contributor note in the README not to open PRs that change legacy
output.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Release 9.1.0: modern uppercase hex, truncation and validation fixes

Bump to 9.1.0. Modern-only: decode uppercase &#X..; hex references
(Rupayon Haldar, #195); preserve fitting post-replacement output during
truncation (emme1t, #193); relocate up-front argument type validation to
the modern path (Jon Bailey, #196). Fix add_uppercase_char atomicity
(Cristian Ramirez, #194). Legacy output unchanged.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Fix release checks for 9.1.0 and tests/ layout

Update tools/check_dist.py to assert version 9.1.0 and the tests/ sdist
layout (tests/test_legacy.py, tests/test_release.py). Add coverage for
reachable branches in the split modules via the public API, restoring the
97% coverage gate without changing legacy behavior.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants