Skip to content

Enhance GIL management and error handling in resampling - #31

Open
shauneccles wants to merge 6 commits into
tuxu:masterfrom
LedFx:perf/more_gil_goodness
Open

shauneccles wants to merge 6 commits into
tuxu:masterfrom
LedFx:perf/more_gil_goodness

Conversation

@shauneccles

@shauneccles shauneccles commented Nov 19, 2025 •

Copy link
Copy Markdown

Release the GIL while resampling, and make Resampler.process() / resample() return all of their output.

Rebased on current master (which already has #30 and #36). Includes the tests from #32, as suggested in review.

GIL (Tino's commit from #14): src_process, src_simple and src_callback_read run with the GIL released, so resampling scales across threads and run_in_executor. The callback takes the GIL for its whole body: the buffer_info destructor also releases a Python buffer. Exceptions raised in the callback (validation errors or the Python callback's own) are stored and re-raised by read() instead of unwinding through libsamplerate's C frames.

Output size: a src_process() call can generate more than ceil(n * ratio):

  • with end_of_input, the converter flushes the input it held back, about ratio × its filter half-length (144 input frames for sinc_best, 20 for sinc_fastest, 0 for linear), so up to ~37k frames at ratio 256;
  • after a ratio change libsamplerate ramps from the previous ratio, which has no fixed bound.

On master that output is silently dropped (e.g. process(x, 4.0) then process(200_000 frames, 0.25, end_of_input=True) returns exactly 50000 frames). A fixed extra allowance (#22) only moves the limit. process() now keeps calling src_process() with a larger buffer until libsamplerate stops short of filling it; resample() goes through the same path, since src_simple() is src_new() + one src_process() + src_delete() and can't continue. This supersedes #22.

Thread safety: releasing the GIL means two Python threads can drive one resampler at once. Several threads sharing a Resampler corrupted its state (Bad length in prepare_data ()), and __exit__ from another thread during read() segfaulted. Each resampler now raises RuntimeError on concurrent or re-entrant use; separate objects still run in parallel.

CallbackResampler.clone() (existing bugs in code this PR touches): src_clone copies the callback pointer, so a clone called back into the original and segfaulted once the original was freed; it also read the original's freed input buffer. Both fixed.

Tests: threading/asyncio performance tests from #32 (speedups reported as warnings, not failures, since CI runners vary), plus regression tests for the high-ratio flush, the ratio change, and an exception raised in the callback. The first two fail on master. CI now installs pytest-asyncio.

@shauneccles

shauneccles commented Nov 19, 2025 •

Copy link
Copy Markdown
Author

For reference this is the results of the asyncio and threading test on my local machine.

PS C:\Users\shaun\python-samplerate-ledfx> uv run pytest tests/test_asyncio_performance.py -s
Uninstalled 1 package in 8ms
Installed 1 package in 10ms
================================================================ test session starts =================================================================
platform win32 -- Python 3.11.9, pytest-9.0.1, pluggy-1.6.0
rootdir: C:\Users\shaun\python-samplerate-ledfx
configfile: pyproject.toml
plugins: asyncio-1.3.0
asyncio: mode=Mode.STRICT, debug=False, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
collected 28 items                                                                                                                                    

tests\test_asyncio_performance.py
default loop - sinc_fastest async with ThreadPoolExecutor (2 concurrent):
  Sequential: 0.0501s
  Parallel: 0.0259s
  Speedup: 1.93x
  Platform: AMD64
  ✓ Performance meets expectations (1.1x)
.
default loop - sinc_fastest async with ThreadPoolExecutor (4 concurrent):
  Sequential: 0.0993s
  Parallel: 0.0259s
  Speedup: 3.84x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
default loop - sinc_fastest async with ThreadPoolExecutor (8 concurrent):
  Sequential: 0.1991s
  Parallel: 0.0302s
  Speedup: 6.60x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
default loop - sinc_medium async with ThreadPoolExecutor (2 concurrent):
  Sequential: 0.0987s
  Parallel: 0.0495s
  Speedup: 1.99x
  Platform: AMD64
  ✓ Performance meets expectations (1.1x)
.
default loop - sinc_medium async with ThreadPoolExecutor (4 concurrent):
  Sequential: 0.1959s
  Parallel: 0.0504s
  Speedup: 3.89x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
default loop - sinc_medium async with ThreadPoolExecutor (8 concurrent):
  Sequential: 0.3948s
  Parallel: 0.0660s
  Speedup: 5.98x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
default loop - sinc_best async with ThreadPoolExecutor (2 concurrent):
  Sequential: 0.3053s
  Parallel: 0.1604s
  Speedup: 1.90x
  Platform: AMD64
  ✓ Performance meets expectations (1.1x)
.
default loop - sinc_best async with ThreadPoolExecutor (4 concurrent):
  Sequential: 0.6138s
  Parallel: 0.1547s
  Speedup: 3.97x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
default loop - sinc_best async with ThreadPoolExecutor (8 concurrent):
  Sequential: 1.2228s
  Parallel: 0.1897s
  Speedup: 6.44x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
winloop loop - sinc_fastest async with ThreadPoolExecutor (2 concurrent):
  Sequential: 0.0499s
  Parallel: 0.0257s
  Speedup: 1.94x
  Platform: AMD64
  ✓ Performance meets expectations (1.1x)
.
winloop loop - sinc_fastest async with ThreadPoolExecutor (4 concurrent):
  Sequential: 0.1001s
  Parallel: 0.0392s
  Speedup: 2.55x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
winloop loop - sinc_fastest async with ThreadPoolExecutor (8 concurrent):
  Sequential: 0.1988s
  Parallel: 0.0373s
  Speedup: 5.33x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
winloop loop - sinc_medium async with ThreadPoolExecutor (2 concurrent):
  Sequential: 0.0979s
  Parallel: 0.0495s
  Speedup: 1.98x
  Platform: AMD64
  ✓ Performance meets expectations (1.1x)
.
winloop loop - sinc_medium async with ThreadPoolExecutor (4 concurrent):
  Sequential: 0.1960s
  Parallel: 0.0576s
  Speedup: 3.40x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
winloop loop - sinc_medium async with ThreadPoolExecutor (8 concurrent):
  Sequential: 0.3957s
  Parallel: 0.0548s
  Speedup: 7.22x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
winloop loop - sinc_best async with ThreadPoolExecutor (2 concurrent):
  Sequential: 0.3051s
  Parallel: 0.1530s
  Speedup: 1.99x
  Platform: AMD64
  ✓ Performance meets expectations (1.1x)
.
winloop loop - sinc_best async with ThreadPoolExecutor (4 concurrent):
  Sequential: 0.6071s
  Parallel: 0.1605s
  Speedup: 3.78x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
winloop loop - sinc_best async with ThreadPoolExecutor (8 concurrent):
  Sequential: 1.2146s
  Parallel: 0.1851s
  Speedup: 6.56x
  Platform: AMD64
  ✓ Performance meets expectations (1.2x)
.
default loop - sinc_fastest blocking vs executor:
  Without executor (blocks loop): 0.0100s
  With ThreadPoolExecutor: 0.0057s
  Improvement: 1.77x
  ✓ Executor performance meets expectations
.
winloop loop - sinc_fastest blocking vs executor:
  Without executor (blocks loop): 0.0100s
  With ThreadPoolExecutor: 0.0056s
  Improvement: 1.80x
  ✓ Executor performance meets expectations
.
default loop - 2 concurrent tasks - ThreadPool vs ProcessPool:
  ThreadPoolExecutor: 0.0106s
  ProcessPoolExecutor: 0.1737s
  Ratio: 16.39x
  → ThreadPool is faster
    (GIL release makes ThreadPool competitive with ProcessPool)
.
default loop - 4 concurrent tasks - ThreadPool vs ProcessPool:
  ThreadPoolExecutor: 0.0109s
  ProcessPoolExecutor: 0.1873s
  Ratio: 17.14x
  → ThreadPool is faster
    (GIL release makes ThreadPool competitive with ProcessPool)
.
winloop loop - 2 concurrent tasks - ThreadPool vs ProcessPool:
  ThreadPoolExecutor: 0.0105s
  ProcessPoolExecutor: 0.1676s
  Ratio: 16.04x
  → ThreadPool is faster
    (GIL release makes ThreadPool competitive with ProcessPool)
.
winloop loop - 4 concurrent tasks - ThreadPool vs ProcessPool:
  ThreadPoolExecutor: 0.0110s
  ProcessPoolExecutor: 0.1850s
  Ratio: 16.89x
  → ThreadPool is faster
    (GIL release makes ThreadPool competitive with ProcessPool)
.
default loop - Mixed I/O and CPU workload:
  Total time: 0.1966s
  Tasks completed: 5
  ✓ Performance meets expectations (< 0.35s)
.
winloop loop - Mixed I/O and CPU workload:
  Total time: 0.2021s
  Tasks completed: 5
  ✓ Performance meets expectations (< 0.35s)
.
======================================================================
Asyncio Performance Report
======================================================================

Test Configuration:
  Sample rate: 44100 Hz
  Duration: 5.0 seconds (220500 samples)
  Conversion ratio: 2.0x
  Executor: ThreadPoolExecutor

----------------------------------------------------------------------
Converter: sinc_fastest
----------------------------------------------------------------------
  1 concurrent task (baseline):
    Execution time: 0.0253s
  2 concurrent tasks:
    Parallel execution time: 0.0254s
    Equivalent sequential time: 0.0507s (2 × 0.0253s)
    Speedup: 2.00x
    Parallel efficiency: 99.9%
  4 concurrent tasks:
    Parallel execution time: 0.0265s
    Equivalent sequential time: 0.1013s (4 × 0.0253s)
    Speedup: 3.82x
    Parallel efficiency: 95.6%

----------------------------------------------------------------------
Converter: sinc_medium
----------------------------------------------------------------------
  1 concurrent task (baseline):
    Execution time: 0.0493s
  2 concurrent tasks:
    Parallel execution time: 0.0497s
    Equivalent sequential time: 0.0985s (2 × 0.0493s)
    Speedup: 1.98x
    Parallel efficiency: 99.2%
  4 concurrent tasks:
    Parallel execution time: 0.0506s
    Equivalent sequential time: 0.1970s (4 × 0.0493s)
    Speedup: 3.90x
    Parallel efficiency: 97.4%

----------------------------------------------------------------------
Converter: sinc_best
----------------------------------------------------------------------
  1 concurrent task (baseline):
    Execution time: 0.1524s
  2 concurrent tasks:
    Parallel execution time: 0.1530s
    Equivalent sequential time: 0.3049s (2 × 0.1524s)
    Speedup: 1.99x
    Parallel efficiency: 99.6%
  4 concurrent tasks:
    Parallel execution time: 0.1579s
    Equivalent sequential time: 0.6098s (4 × 0.1524s)
    Speedup: 3.86x
    Parallel efficiency: 96.5%
.
======================================================================
Asyncio Performance Report
======================================================================

Test Configuration:
  Sample rate: 44100 Hz
  Duration: 5.0 seconds (220500 samples)
  Conversion ratio: 2.0x
  Executor: ThreadPoolExecutor

----------------------------------------------------------------------
Converter: sinc_fastest
----------------------------------------------------------------------
  1 concurrent task (baseline):
    Execution time: 0.0253s
  2 concurrent tasks:
    Parallel execution time: 0.0290s
    Equivalent sequential time: 0.0505s (2 × 0.0253s)
    Speedup: 1.74x
    Parallel efficiency: 87.2%
  4 concurrent tasks:
    Parallel execution time: 0.0264s
    Equivalent sequential time: 0.1010s (4 × 0.0253s)
    Speedup: 3.83x
    Parallel efficiency: 95.9%

----------------------------------------------------------------------
Converter: sinc_medium
----------------------------------------------------------------------
  1 concurrent task (baseline):
    Execution time: 0.0497s
  2 concurrent tasks:
    Parallel execution time: 0.0579s
    Equivalent sequential time: 0.0995s (2 × 0.0497s)
    Speedup: 1.72x
    Parallel efficiency: 85.9%
  4 concurrent tasks:
    Parallel execution time: 0.0609s
    Equivalent sequential time: 0.1990s (4 × 0.0497s)
    Speedup: 3.27x
    Parallel efficiency: 81.7%

----------------------------------------------------------------------
Converter: sinc_best
----------------------------------------------------------------------
  1 concurrent task (baseline):
    Execution time: 0.1542s
  2 concurrent tasks:
    Parallel execution time: 0.1538s
    Equivalent sequential time: 0.3083s (2 × 0.1542s)
    Speedup: 2.00x
    Parallel efficiency: 100.2%
  4 concurrent tasks:
    Parallel execution time: 0.1596s
    Equivalent sequential time: 0.6167s (4 × 0.1542s)
    Speedup: 3.86x
    Parallel efficiency: 96.6%
PS C:\Users\shaun\python-samplerate-ledfx> uv run pytest .\tests\test_threading_performance.py -s
================================================================ test session starts =================================================================
platform win32 -- Python 3.11.9, pytest-9.0.1, pluggy-1.6.0
rootdir: C:\Users\shaun\python-samplerate-ledfx
configfile: pyproject.toml
plugins: asyncio-1.3.0
asyncio: mode=Mode.STRICT, debug=False, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
collected 38 items                                                                                                                                    

tests\test_threading_performance.py
sinc_fastest with 2 threads:
  Sequential: 0.0500s
  Parallel: 0.0259s
  Speedup: 1.93x
  Platform: AMD64
  Individual thread times: ['0.0248s', '0.0248s']
  ✓ Performance meets expectations (1.2x)
.
sinc_fastest with 4 threads:
  Sequential: 0.0995s
  Parallel: 0.0258s
  Speedup: 3.86x
  Platform: AMD64
  Individual thread times: ['0.0251s', '0.0249s', '0.0249s', '0.0249s']
  ✓ Performance meets expectations (1.35x)
.
sinc_fastest with 6 threads:
  Sequential: 0.1492s
  Parallel: 0.0364s
  Speedup: 4.10x
  Platform: AMD64
  Individual thread times: ['0.0318s', '0.0293s', '0.0356s', '0.0252s', '0.0253s', '0.0250s']
  ✓ Performance meets expectations (1.35x)
.
sinc_fastest with 8 threads:
  Sequential: 0.2029s
  Parallel: 0.0353s
  Speedup: 5.74x
  Platform: AMD64
  Individual thread times: ['0.0248s', '0.0248s', '0.0248s', '0.0278s', '0.0344s', '0.0318s', '0.0249s', '0.0248s']
  ✓ Performance meets expectations (1.35x)
.
sinc_medium with 2 threads:
  Sequential: 0.0979s
  Parallel: 0.0576s
  Speedup: 1.70x
  Platform: AMD64
  Individual thread times: ['0.0570s', '0.0572s']
  ✓ Performance meets expectations (1.2x)
.
sinc_medium with 4 threads:
  Sequential: 0.1957s
  Parallel: 0.0590s
  Speedup: 3.32x
  Platform: AMD64
  Individual thread times: ['0.0576s', '0.0584s', '0.0506s', '0.0500s']
  ✓ Performance meets expectations (1.35x)
.
sinc_medium with 6 threads:
  Sequential: 0.2950s
  Parallel: 0.0629s
  Speedup: 4.69x
  Platform: AMD64
  Individual thread times: ['0.0614s', '0.0624s', '0.0508s', '0.0503s', '0.0518s', '0.0494s']
  ✓ Performance meets expectations (1.35x)
.
sinc_medium with 8 threads:
  Sequential: 0.3993s
  Parallel: 0.0635s
  Speedup: 6.29x
  Platform: AMD64
  Individual thread times: ['0.0588s', '0.0593s', '0.0628s', '0.0551s', '0.0502s', '0.0603s', '0.0596s', '0.0616s']
  ✓ Performance meets expectations (1.35x)
.
sinc_best with 2 threads:
  Sequential: 0.3043s
  Parallel: 0.1603s
  Speedup: 1.90x
  Platform: AMD64
  Individual thread times: ['0.1600s', '0.1583s']
  ✓ Performance meets expectations (1.2x)
.
sinc_best with 4 threads:
  Sequential: 0.6091s
  Parallel: 0.1609s
  Speedup: 3.78x
  Platform: AMD64
  Individual thread times: ['0.1605s', '0.1572s', '0.1548s', '0.1527s']
  ✓ Performance meets expectations (1.35x)
.
sinc_best with 6 threads:
  Sequential: 0.9121s
  Parallel: 0.1678s
  Speedup: 5.43x
  Platform: AMD64
  Individual thread times: ['0.1520s', '0.1667s', '0.1619s', '0.1626s', '0.1535s', '0.1667s']
  ✓ Performance meets expectations (1.35x)
.
sinc_best with 8 threads:
  Sequential: 1.2174s
  Parallel: 0.1758s
  Speedup: 6.93x
  Platform: AMD64
  Individual thread times: ['0.1559s', '0.1645s', '0.1569s', '0.1597s', '0.1551s', '0.1746s', '0.1643s', '0.1612s']
  ✓ Performance meets expectations (1.35x)
.
sinc_fastest Resampler.process() with 2 threads:
  Sequential: 0.0499s
  Parallel: 0.0254s
  Speedup: 1.96x
  Platform: AMD64
  Individual thread times: ['0.0249s', '0.0249s']
  ✓ Performance meets expectations (1.1x)
.
sinc_fastest Resampler.process() with 4 threads:
  Sequential: 0.1001s
  Parallel: 0.0264s
  Speedup: 3.79x
  Platform: AMD64
  Individual thread times: ['0.0252s', '0.0251s', '0.0257s', '0.0252s']
  ✓ Performance meets expectations (1.25x)
.
sinc_fastest Resampler.process() with 6 threads:
  Sequential: 0.1502s
  Parallel: 0.0323s
  Speedup: 4.65x
  Platform: AMD64
  Individual thread times: ['0.0249s', '0.0311s', '0.0249s', '0.0250s', '0.0313s', '0.0249s']
  ✓ Performance meets expectations (1.25x)
.
sinc_fastest Resampler.process() with 8 threads:
  Sequential: 0.2012s
  Parallel: 0.0271s
  Speedup: 7.41x
  Platform: AMD64
  Individual thread times: ['0.0248s', '0.0251s', '0.0255s', '0.0249s', '0.0250s', '0.0250s', '0.0254s', '0.0253s']
  ✓ Performance meets expectations (1.25x)
.
sinc_medium Resampler.process() with 2 threads:
  Sequential: 0.0981s
  Parallel: 0.0495s
  Speedup: 1.98x
  Platform: AMD64
  Individual thread times: ['0.0489s', '0.0491s']
  ✓ Performance meets expectations (1.1x)
.
sinc_medium Resampler.process() with 4 threads:
  Sequential: 0.1961s
  Parallel: 0.0506s
  Speedup: 3.88x
  Platform: AMD64
  Individual thread times: ['0.0488s', '0.0489s', '0.0489s', '0.0497s']
  ✓ Performance meets expectations (1.25x)
.
sinc_medium Resampler.process() with 6 threads:
  Sequential: 0.2939s
  Parallel: 0.0510s
  Speedup: 5.77x
  Platform: AMD64
  Individual thread times: ['0.0489s', '0.0489s', '0.0488s', '0.0492s', '0.0497s', '0.0490s']
  ✓ Performance meets expectations (1.25x)
.
sinc_medium Resampler.process() with 8 threads:
  Sequential: 0.3920s
  Parallel: 0.0575s
  Speedup: 6.81x
  Platform: AMD64
  Individual thread times: ['0.0494s', '0.0491s', '0.0493s', '0.0494s', '0.0499s', '0.0558s', '0.0498s', '0.0558s']
  ✓ Performance meets expectations (1.25x)
.
sinc_best Resampler.process() with 2 threads:
  Sequential: 0.3049s
  Parallel: 0.1525s
  Speedup: 2.00x
  Platform: AMD64
  Individual thread times: ['0.1521s', '0.1519s']
  ✓ Performance meets expectations (1.1x)
.
sinc_best Resampler.process() with 4 threads:
  Sequential: 0.6109s
  Parallel: 0.1683s
  Speedup: 3.63x
  Platform: AMD64
  Individual thread times: ['0.1580s', '0.1537s', '0.1675s', '0.1623s']
  ✓ Performance meets expectations (1.25x)
.
sinc_best Resampler.process() with 6 threads:
  Sequential: 0.9163s
  Parallel: 0.1729s
  Speedup: 5.30x
  Platform: AMD64
  Individual thread times: ['0.1599s', '0.1709s', '0.1721s', '0.1651s', '0.1653s', '0.1574s']
  ✓ Performance meets expectations (1.25x)
.
sinc_best Resampler.process() with 8 threads:
  Sequential: 1.2241s
  Parallel: 0.1809s
  Speedup: 6.77x
  Platform: AMD64
  Individual thread times: ['0.1600s', '0.1637s', '0.1769s', '0.1727s', '0.1661s', '0.1792s', '0.1726s', '0.1598s']
  ✓ Performance meets expectations (1.25x)
.
sinc_fastest CallbackResampler with 2 threads:
  Sequential: 0.0504s
  Parallel: 0.0258s
  Speedup: 1.95x
  Platform: AMD64
  Individual thread times: ['0.0251s', '0.0252s']
  ✓ Performance meets expectations (1.2x)
.
sinc_fastest CallbackResampler with 4 threads:
  Sequential: 0.1012s
  Parallel: 0.0267s
  Speedup: 3.79x
  Platform: AMD64
  Individual thread times: ['0.0258s', '0.0254s', '0.0258s', '0.0258s']
  ✓ Performance meets expectations (1.2x)
.
sinc_fastest CallbackResampler with 6 threads:
  Sequential: 0.1523s
  Parallel: 0.0273s
  Speedup: 5.58x
  Platform: AMD64
  Individual thread times: ['0.0270s', '0.0252s', '0.0257s', '0.0256s', '0.0256s', '0.0258s']
  ✓ Performance meets expectations (1.2x)
.
sinc_fastest CallbackResampler with 8 threads:
  Sequential: 0.2013s
  Parallel: 0.0404s
  Speedup: 4.99x
  Platform: AMD64
  Individual thread times: ['0.0261s', '0.0397s', '0.0397s', '0.0261s', '0.0252s', '0.0255s', '0.0256s', '0.0260s']
  ✓ Performance meets expectations (1.2x)
.
sinc_medium CallbackResampler with 2 threads:
  Sequential: 0.0985s
  Parallel: 0.0499s
  Speedup: 1.97x
  Platform: AMD64
  Individual thread times: ['0.0491s', '0.0495s']
  ✓ Performance meets expectations (1.2x)
.
sinc_medium CallbackResampler with 4 threads:
  Sequential: 0.1970s
  Parallel: 0.0593s
  Speedup: 3.32x
  Platform: AMD64
  Individual thread times: ['0.0585s', '0.0588s', '0.0513s', '0.0495s']
  ✓ Performance meets expectations (1.2x)
.
sinc_medium CallbackResampler with 6 threads:
  Sequential: 0.2957s
  Parallel: 0.0554s
  Speedup: 5.34x
  Platform: AMD64
  Individual thread times: ['0.0539s', '0.0495s', '0.0495s', '0.0495s', '0.0541s', '0.0495s']
  ✓ Performance meets expectations (1.2x)
.
sinc_medium CallbackResampler with 8 threads:
  Sequential: 0.3935s
  Parallel: 0.0775s
  Speedup: 5.08x
  Platform: AMD64
  Individual thread times: ['0.0772s', '0.0696s', '0.0517s', '0.0592s', '0.0596s', '0.0503s', '0.0567s', '0.0668s']
  ✓ Performance meets expectations (1.2x)
.
sinc_best CallbackResampler with 2 threads:
  Sequential: 0.3054s
  Parallel: 0.1529s
  Speedup: 2.00x
  Platform: AMD64
  Individual thread times: ['0.1525s', '0.1521s']
  ✓ Performance meets expectations (1.2x)
.
sinc_best CallbackResampler with 4 threads:
  Sequential: 0.6106s
  Parallel: 0.1574s
  Speedup: 3.88x
  Platform: AMD64
  Individual thread times: ['0.1570s', '0.1566s', '0.1561s', '0.1524s']
  ✓ Performance meets expectations (1.2x)
.
sinc_best CallbackResampler with 6 threads:
  Sequential: 0.9167s
  Parallel: 0.1798s
  Speedup: 5.10x
  Platform: AMD64
  Individual thread times: ['0.1712s', '0.1751s', '0.1734s', '0.1787s', '0.1548s', '0.1551s']
  ✓ Performance meets expectations (1.2x)
.
sinc_best CallbackResampler with 8 threads:
  Sequential: 1.2261s
  Parallel: 0.1708s
  Speedup: 7.18x
  Platform: AMD64
  Individual thread times: ['0.1645s', '0.1572s', '0.1671s', '0.1557s', '0.1639s', '0.1564s', '0.1682s', '0.1685s']
  ✓ Performance meets expectations (1.2x)
..
======================================================================
GIL Release Performance Report
======================================================================

Test Configuration:
  Sample rate: 44100 Hz
  Duration: 5.0 seconds (220500 samples)
  Conversion ratio: 2.0x

----------------------------------------------------------------------
Converter: sinc_fastest
----------------------------------------------------------------------
  1 thread (baseline):
    Execution time: 0.0249s
  2 threads (parallel):
    Parallel execution time: 0.0256s
    Equivalent sequential time: 0.0498s (2 × 0.0249s)
    Speedup: 1.95x
    Parallel efficiency: 97.3%
    Avg thread time: 0.0250s
  4 threads (parallel):
    Parallel execution time: 0.0257s
    Equivalent sequential time: 0.0995s (4 × 0.0249s)
    Speedup: 3.87x
    Parallel efficiency: 96.8%
    Avg thread time: 0.0249s

----------------------------------------------------------------------
Converter: sinc_medium
----------------------------------------------------------------------
  1 thread (baseline):
    Execution time: 0.0491s
  2 threads (parallel):
    Parallel execution time: 0.0494s
    Equivalent sequential time: 0.0981s (2 × 0.0491s)
    Speedup: 1.98x
    Parallel efficiency: 99.2%
    Avg thread time: 0.0489s
  4 threads (parallel):
    Parallel execution time: 0.0505s
    Equivalent sequential time: 0.1962s (4 × 0.0491s)
    Speedup: 3.89x
    Parallel efficiency: 97.2%
    Avg thread time: 0.0496s

----------------------------------------------------------------------
Converter: sinc_best
----------------------------------------------------------------------
  1 thread (baseline):
    Execution time: 0.1527s
  2 threads (parallel):
    Parallel execution time: 0.1534s
    Equivalent sequential time: 0.3055s (2 × 0.1527s)
    Speedup: 1.99x
    Parallel efficiency: 99.6%
    Avg thread time: 0.1529s
  4 threads (parallel):
    Parallel execution time: 0.1532s
    Equivalent sequential time: 0.6110s (4 × 0.1527s)
    Speedup: 3.99x
    Parallel efficiency: 99.7%
    Avg thread time: 0.1524s

Looking at #32 which has same tests on same machine at (about) the same time the headlines from this PR is:

  • Releasing the GIL is a very good thing
    • Basically a linear performance improvement with more threads/concurrency
  • Asyncio + ThreadPoolExecutor is probably the way to go for concurrent resampling

@fakufaku fakufaku left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks pretty good to me. Please just sync with master.
Would it make sense to combine this with your other PR that contains the tests?

Comment thread .github/workflows/pythonpackage.yml Outdated
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: [3.8, 3.9, "3.10", "3.11", "3.12"]
python-version: [3.9, "3.10", "3.11", "3.12", "3.13", "3.14"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This has been merged already. Could you please sync with master?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rebased on master, so this is no longer in the diff.

Comment thread src/samplerate.cpp Outdated
#endif

// This value was empirically and somewhat arbitrarily chosen; increase it for further safety.
#define END_OF_INPUT_EXTRA_OUTPUT_FRAMES 10000

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please add some more details about what this constant controls?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's now OUTPUT_HEADROOM_FRAMES (1024), with a comment: the output frames allocated on top of ceil(input_frames * ratio). Since process() now grows the buffer when libsamplerate fills it, the value no longer affects correctness, only how often a regrowth happens.

Comment thread src/samplerate.cpp Outdated
// be more than the expected number of output samples during mid-stream
// steady-state processing. (Also, when the stream is started, the number
// of output samples generated will generally be zero or otherwise less
// than the number of samples in mid-stream processing.)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a maximum we could expect?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For a fixed ratio, yes: the end_of_input flush emits roughly ratio × the input the converter holds back, i.e. its filter half-length. Measured: 144 input frames for sinc_best (tail of 2304 at ratio 16, 9216 at ratio 64), 20 for sinc_fastest, 0 for linear. That's up to ~37k frames at ratio 256, so the 10000 constant fails with sinc_best above ratio ~70.

After a ratio change, though, there's no fixed maximum: libsamplerate ramps from the previous ratio, so process(x, 4.0) followed by process(200_000 frames, 0.25) produces far more than 50000. So I dropped the idea of a maximum; see the next comment.

Comment thread src/samplerate.cpp Outdated
output.resize(out_shape);
} else if ((size_t)output_frames_gen >= new_size) {
// This means our fudge factor is too small.
throw std::runtime_error("Generated more output samples than expected!");

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rather than failing, is it possible to truncate the output?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Truncating would lose audio. When the output buffer fills, libsamplerate hasn't consumed all the input yet, so whatever we cut is gone (master does this silently today). Instead process() now keeps calling src_process() on the remaining input with a larger buffer until libsamplerate stops short of filling it, and resample() uses the same path. test_process_flush_at_high_ratio and test_process_after_ratio_change cover both cases and fail on master.

tuxu and others added 5 commits September 27, 2026 13:59
… use

Output size: a src_process() call can generate more than
ceil(input_frames * ratio). With end_of_input the converter flushes the
input it held back (about ratio x its filter half-length: 144 input
frames for sinc_best, so ~37k output frames at ratio 256), and after a
ratio change libsamplerate ramps from the previous ratio, which has no
fixed bound. Today process() silently drops that output, and a fixed
extra allowance (#22) only moves the failure. Resampler.process() now
keeps calling src_process() with a larger buffer until libsamplerate
stops short of filling it. resample() uses the same path, since
src_simple() is src_new() + one src_process() + src_delete() and cannot
continue.

Callback: hold the GIL for the whole of the_callback_func, since the
buffer_info destructor releases a Python buffer, and store any exception
(ours or one raised by the Python callback) instead of letting it unwind
through libsamplerate's C frames; read() re-raises it.

Also run the asyncio tests in CI (pytest-asyncio).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Releasing the GIL lets two Python threads drive one SRC_STATE at once:
four threads calling process() on one Resampler got "Internal error:
Bad length in prepare_data ()", and __exit__ from another thread during
read() segfaulted. Every method that touches the state now takes a
per-object in-use flag and raises RuntimeError if it's already set. A
mutex could deadlock, since the holder needs the GIL to grow the output
buffer while a waiter may hold it.

Two existing clone bugs, in the same code:
- src_clone() copies the callback data pointer, so a clone called back
  into the original CallbackResampler and segfaulted once the original
  was freed. read() now publishes the active resampler in a
  thread_local, which the callback uses instead.
- A clone's libsamplerate state still points into the original's last
  callback buffer; the copy now holds a reference to it, where it used
  to read freed memory.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@shauneccles

Copy link
Copy Markdown
Author

One more commit after reviewing the GIL handling again: releasing the GIL let two threads use one resampler at the same time (corrupted state, and a segfault via __exit__), so each resampler now raises RuntimeError on concurrent use. It also fixes two existing CallbackResampler.clone() bugs: the clone called back into the original, and read its freed input buffer. I reproduced each one before fixing it; there are regression tests for the deterministic ones. Details are in the description.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants