Skip to content

feat: do not entirely disable connection pooling for periodic connections - #2440

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 2 commits into
mainfrom
yannham/low-timeout-connection-pooling
Sep 2, 2026
Merged

feat: do not entirely disable connection pooling for periodic connections#2440
gh-worker-dd-mergequeue-cf854d[bot] merged 2 commits into
mainfrom
yannham/low-timeout-connection-pooling

Conversation

@yannham

@yannham yannham commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

What does this PR do?

This PR slightly changes the HTTP behavior for periodic connections (most flushes to the agent/agentless intake). Instead of disabling connection pooling to avoid data races entirely, it restricts the lifetime of pooled connections.

This PR aims to be a backward-compatible hotfix. A subsequent PR is coming with renaming (since no_connection_pooling isn't really true anymore) and mirroring the change in libdd-http-client and libdd-agent-client as well.

Motivation

Mitigates APMS-20441 / DataDog/dd-trace-py#19915: some telemetry events can generate several requests for short-lived scripts. In the agentless case, this means opening several new HTTPS connection in a row, which is slow.

Additional Notes

Applying the setting to all periodic connections sounds reasonable, as multiple requests could reasonably be issued for other things than telemetry. It should impact single-request workflows.

How to test the change?

This was tested with the repro in 20441, reducing the shutdown delay to 0.750ms locally, which indicates there's indeed a single HTTPS connection (vs double before the change).

@yannham
yannham requested review from a team as code owners September 1, 2026 12:57
@yannham
yannham requested review from danyal002 and vjfridge and removed request for a team September 1, 2026 12:57
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 1, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-09-01T13:01:50.372159Z 8d26774 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@datadog-official datadog-official Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Datadog Autotest: PASS

More details

Both HTTP backends set the same five-second idle limit for periodic clients. The changed call sites use the new periodic policy consistently.

Was this helpful? React 👍 or 👎

Open Bits AI session

🤖 Datadog Autotest · Commit 8d26774 · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

@datadog-official

datadog-official Bot commented Sep 1, 2026

Copy link
Copy Markdown

Tests

All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
Patch Coverage: 100.00%
Overall Coverage: 76.88% (-0.06%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 1444009 | Docs | View more details | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Artifact Size Benchmark Report

aarch64-alpine-linux-musl
Artifact Baseline Commit Change
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.a 90.77 MB 90.77 MB +0% (+1.48 KB) 👌
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.so 8.39 MB 8.39 MB 0% (0 B) 👌
aarch64-unknown-linux-gnu
Artifact Baseline Commit Change
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.so 11.29 MB 11.29 MB +0% (+192 B) 👌
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.a 101.99 MB 101.99 MB +0% (+816 B) 👌
libdatadog-x64-windows
Artifact Baseline Commit Change
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.dll 27.03 MB 27.03 MB +0% (+512 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.lib 96.08 KB 96.08 KB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.pdb 181.69 MB 181.69 MB 0% (0 B) 👌
/libdatadog-x64-windows/debug/static/datadog_profiling_ffi.lib 766.69 MB 765.98 MB --.09% (-727.71 KB) 💪
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.dll 8.91 MB 8.91 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.lib 96.08 KB 96.08 KB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.pdb 26.05 MB 26.05 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/static/datadog_profiling_ffi.lib 51.84 MB 51.84 MB +0% (+1.10 KB) 👌
libdatadog-x86-windows
Artifact Baseline Commit Change
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.dll 23.57 MB 23.57 MB +0% (+512 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.lib 97.58 KB 97.58 KB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.pdb 186.76 MB 186.77 MB +0% (+8.00 KB) 👌
/libdatadog-x86-windows/debug/static/datadog_profiling_ffi.lib 754.57 MB 755.23 MB +.08% (+674.09 KB) 🔍
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.dll 6.88 MB 6.88 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.lib 97.58 KB 97.58 KB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.pdb 28.00 MB 28.00 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/static/datadog_profiling_ffi.lib 49.34 MB 49.34 MB +0% (+1.31 KB) 👌
x86_64-alpine-linux-musl
Artifact Baseline Commit Change
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.a 80.93 MB 80.94 MB +0% (+888 B) 👌
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.so 9.33 MB 9.33 MB 0% (0 B) 👌
x86_64-unknown-linux-gnu
Artifact Baseline Commit Change
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.a 96.69 MB 96.69 MB +0% (+872 B) 👌
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.so 11.37 MB 11.37 MB +0% (+144 B) 👌

Periodic clients (telemetry flushes, remote config polling, etc.) used to
disable connection pooling entirely, which meant every request paid a
full TLS handshake. That is quite costly for agentless traffic (on the
order of magnitude of 0.5s per connection), and short-lived apps send
several separate requests in a short span of time.

Instead of disabling pooling, `new_client_periodic` now pools
connections with a small idle timeout (5s), much smaller than typical
keep-alive timeouts on the receiving end, so we still avoid reusing a
connection the receiver may have closed.
@pr-commenter

pr-commenter Bot commented Sep 1, 2026

Copy link
Copy Markdown

Benchmarks

Comparison

Benchmark execution time: 2026-09-02 09:09:36

Comparing candidate commit 1444009 in PR branch yannham/low-timeout-connection-pooling with baseline commit 1887a57 in branch main.

Found 0 performance improvements and 0 performance regressions! Performance is the same for 135 metrics, 1 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

Benchmark execution time: 2026-09-02 09:09:59

Comparing candidate commit 1444009 in PR branch yannham/low-timeout-connection-pooling with baseline commit 1887a57 in branch main.

Found 5 performance improvements and 0 performance regressions! Performance is the same for 103 metrics, 10 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:datadog_sample_span/resource_pattern_rule_matching/wall_time

  • 🟩 execution_time [-16.206ns; -16.084ns] or [-5.975%; -5.930%]

scenario:glob_matcher/ascii_wildcard_question_match/wall_time

  • 🟩 execution_time [-23.387ns; -23.355ns] or [-38.829%; -38.776%]

scenario:glob_matcher/ascii_wildcard_star_match/wall_time

  • 🟩 execution_time [-21.105ns; -21.067ns] or [-36.417%; -36.352%]

scenario:trace_buffer/2_senders/no_delay

  • 🟩 execution_time [-45.519µs; -38.903µs] or [-4.952%; -4.232%]
  • 🟩 throughput [+87249.122op/s; +102011.416op/s] or [+4.454%; +5.208%]

Candidate

Omitted due to size.

Baseline

Omitted due to size.

Comment thread libdd-common/src/http_common.rs
Comment thread libdd-capabilities-impl/src/http.rs
@yannham

yannham commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Sep 2, 2026

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-09-02 08:33:17 UTC ℹ️ Start processing command /merge


2026-09-02 08:33:25 UTC ℹ️ MergeQueue: Pull request is not mergeable yet

It will be processed automatically as soon as GitHub reports it as mergeable. View in MergeQueue UI.

  • Run /code blockers to see what is blocking it.
  • Run /remove to cancel it.

2026-09-02 09:15:38 UTC ℹ️ MergeQueue: merge request added to the queue

The expected merge time in main is approximately 49m (p90).


2026-09-02 09:57:38 UTC ℹ️ MergeQueue: This merge request was merged

@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot merged commit a4df07e into main Sep 2, 2026
125 checks passed
@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot deleted the yannham/low-timeout-connection-pooling branch September 2, 2026 09:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants