Skip to content

Preserve fitting post-replacement output during modern truncation - #193

Merged
un33k merged 1 commit into
un33k:masterfrom
emme1t:fix/modern-truncation-preserve-fitting-output
Sep 18, 2026
Merged

un33k merged 1 commit into
un33k:masterfrom
emme1t:fix/modern-truncation-preserve-fitting-output

Conversation

@emme1t

@emme1t emme1t commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

With the modern algorithm, enabling a length limit can change a slug which already fits that limit:

from slugify import slugify

options = dict(
    algorithm="modern",
    word_boundary=True,
    replacements=[("one", "one---")],
    replacement_stage="post",
)

print(slugify("one two", **options))
# one----two
print(slugify("one two", max_length=100, **options))
# one-two

The README specifies that post replacements remain unfiltered. The word-boundary truncation loop drops empty internal tokens, which collapses intentionally repeated delimiters even when the whole output fits. At an exact length limit, the hard-cut branch can also drop an intentionally added trailing delimiter.

Check the final emitted length before truncating. When the entire output fits, return it with the configured separator mapping intact. The length calculation accounts for empty and multi-character separators without constructing an expanded output just to measure it. This change is confined to the modern helper; the default legacy algorithm and public smart_truncate retain their implementations.

Validation on Windows with Python 3.13.13:

  • The original implementation fails 28 new regression subcases covering repeated/boundary delimiters, exact/spare budgets, empty/multi-character separators, word boundaries, and save_order.
  • Full suite: 105 passed, 104 subtests passed. This includes the unchanged legacy suite and its 2,688-case differential comparison against the frozen reference.
  • Strict Mypy, repository-configured pycodestyle and flake8 commands, and git diff --check all pass.

Dojo follow-up: OpenAI Codex assisted with the investigation, patch, regression tests, and this PR description. The reproductions and checks above were executed locally.

@un33k

un33k commented Sep 18, 2026

Copy link
Copy Markdown
Owner

Confirmed and merging. When the whole modern output already fits max_length, the truncation loop was collapsing intentionally repeated delimiters from unfiltered post-replacements (and could drop a trailing delimiter at an exact limit). Checking the emitted length first and returning the mapped output untouched matches the documented 'post replacements remain unfiltered' contract. Confined to the modern helper; legacy and public smart_truncate are untouched. Thanks, @emme1t.

🚀 Generated with Dojo ⛩️

@un33k
un33k merged commit 548b14f into un33k:master Sep 18, 2026
un33k added a commit that referenced this pull request Sep 18, 2026
Bump to 9.1.0. Modern-only: decode uppercase &#X..; hex references
(Rupayon Haldar, #195); preserve fitting post-replacement output during
truncation (emme1t, #193); relocate up-front argument type validation to
the modern path (Jon Bailey, #196). Fix add_uppercase_char atomicity
(Cristian Ramirez, #194). Legacy output unchanged.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️
un33k added a commit that referenced this pull request Sep 18, 2026
* Split frozen legacy pipeline into slugify/_legacy.py

Move the legacy slug pipeline into a dedicated frozen module and make the
public slugify() a thin dispatcher: algorithm='legacy' (default) calls the
frozen _legacy implementation, algorithm='modern' runs the modern pipeline.
Relocate the existing up-front TypeError validation for bool/non-int
max_length and non-str separator (from #196) into the modern path only, so
legacy output is unchanged while modern keeps rejecting invalid types.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Reorganize tests under tests/ with frozen legacy suite

Move the test suite into tests/: the original upstream legacy suite becomes
the frozen tests/test_legacy.py (contents unchanged), mirroring the _legacy.py
code split, alongside tests/test_release.py and the add_uppercase test. The
modern-only bool/separator validation test moves to the modern suite. Update
pyproject testpaths, MANIFEST.in, and tox commands to the tests/ layout.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Document legacy-frozen policy and split in DOJO.md and README

Record that legacy is architecturally frozen (slugify/_legacy.py and
tests/test_legacy.py) and that all new work targets algorithm='modern'.
Add a contributor note in the README not to open PRs that change legacy
output.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Release 9.1.0: modern uppercase hex, truncation and validation fixes

Bump to 9.1.0. Modern-only: decode uppercase &#X..; hex references
(Rupayon Haldar, #195); preserve fitting post-replacement output during
truncation (emme1t, #193); relocate up-front argument type validation to
the modern path (Jon Bailey, #196). Fix add_uppercase_char atomicity
(Cristian Ramirez, #194). Legacy output unchanged.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️

* Fix release checks for 9.1.0 and tests/ layout

Update tools/check_dist.py to assert version 9.1.0 and the tests/ sdist
layout (tests/test_legacy.py, tests/test_release.py). Add coverage for
reachable branches in the split modules via the public API, restoring the
97% coverage gate without changing legacy behavior.

🚀 Generated with [Dojo](https://heydojo.ai) ⛩️
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants