Conversation
This replaces different repeated code blocks that read pgBufferUsage / pgWalUsage, and may have also been running a timer to measure activity, with the new Instrumentation struct and associated helpers. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Discussion:
…ith INSTR_* macros This encapsulates the ownership of these globals better, and will allow a subsequent refactoring. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Andres Freund <andres@anarazel.de> Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com> Discussion: https://www.postgresql.org/message-id/flat/CAP53PkzZ3UotnRrrnXWAv%3DF4avRq9MQ8zU%2BbxoN9tpovEu6fGQ%40mail.gmail.com#fc7140e8af21e07a90a09d7e76b300c4
This adds regression tests that cover some of the expected behaviour around the buffer statistics reported in EXPLAIN ANALYZE, specifically how they behave in parallel query, nested function calls and abort situations. Testing this is challenging because there can be different sources of buffer activity, so we rely on temporary tables where we can to prove that activity was captured and not lost. This supports a future commit that will rework some of the instrumentation logic that could cause areas covered by these tests to fail. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Discussion:
Previously, in order to determine the buffer/WAL usage of a given code section, we utilized continuously incrementing global counters that get updated when the actual activity (e.g. shared block read) occurred, and then calculated a diff when the code section ended. This resulted in a bottleneck for executor node instrumentation specifically, with the function BufferUsageAccumDiff showing up in profiles and in some cases adding up to 10% overhead to an EXPLAIN (ANALYZE, BUFFERS) run. Instead, introduce a stack-based mechanism, where the actual activity writes into the current stack entry. In the case of executor nodes, this means that each node gets its own stack entry that is pushed at InstrStartNode, and popped at InstrEndNode. Stack entries are zero initialized (avoiding the diff mechanism) and get added to their parent entry when they are finalized, i.e. no more modifications can occur. To correctly handle abort situations, any use of instrumentation stacks must involve either a top-level QueryInstrumentation struct, and its associated InstrQueryStart/InstrQueryStop helpers (which use resource owners to handle aborts), or the Instrumentation struct itself with dedicated PG_TRY/PG_FINALLY calls that ensure the stack is in a consistent state after an abort. In tests, the stack-based instrumentation mechanism reduces the overhead of EXPLAIN (ANALYZE, BUFFERS ON, TIMING OFF) for a large COUNT(*) query from about 50% to 22% on top of the actual runtime. Callers of the plain Instrumentation API (InstrAlloc/InstrInitOptions with InstrStart/InstrStop) keep measuring by diffing the global pgBufferUsage/pgWalUsage counters for now. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com> Reviewed-by: Heikki Linnakangas <hlinnaka@iki.fi> Discussion: https://www.postgresql.org/message-id/flat/CAP53PkxrmpECzVFpeeEEHDGe6u625s%2BYkmVv5-gw3L_NDSfbiA%40mail.gmail.com#cb583a08e8e096aa1f093bb178906173
…Usage The stack-based instrumentation left the plain Instrumentation API (InstrAlloc/InstrInitOptions with InstrStart/InstrStop) measuring by diffing the global pgBufferUsage/pgWalUsage counters, and kept pgBufferUsage updated alongside the stack for that purpose. Convert the remaining users of that API to the stack: VACUUM and ANALYZE logging as well as the planning-time measurement of EXPLAIN and EXECUTE use QueryInstrumentation, which handles aborts through the resource owner, while pg_stat_statements wraps its planner and utility statement timing in PG_TRY/PG_FINALLY and the EXPLAIN destination receiver finalizes its entry into the current stack entry. With that, Instrumentation entries always track WAL/buffer usage through the stack, so remove the diff mode along with the global pgBufferUsage and BufferUsageAccumDiff, and update comments that referred to it. Extensions that read pgBufferUsage directly should use InstrStart and InstrStop around the code section they want to measure instead. The related global pgWalUsage is kept for now due to its use in pgstat to track aggregate WAL activity and heap_page_prune_and_freeze for measuring FPIs. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Discussion:
This simplifies the DSM allocations a bit since we don't need to separately allocate WAL and buffer usage, and allows the easier future addition of a third stack-based struct being discussed. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Discussion:
For most queries, the bulk of the overhead of EXPLAIN ANALYZE happens in ExecProcNodeInstr when starting/stopping instrumentation for that node. Previously each ExecProcNodeInstr would check which instrumentation options are active in the InstrStartNode/InstrStopNode calls, and do the corresponding work (timers, instrumentation stack, etc.). These conditionals being checked for each tuple being emitted add up, and cause non-optimal set of instructions to be generated by the compiler. Because we already have an existing mechanism to specify a function pointer when instrumentation is enabled, we can instead create specialized functions that are tailored to the instrumentation options enabled, and avoid conditionals on subsequent ExecProcNodeInstr calls. This results in the overhead for EXPLAIN (ANALYZE, TIMING OFF, BUFFERS OFF) for a stress test with a large COUNT(*) that does many ExecProcNode calls from ~ 20% on top of actual runtime to ~ 3%. When using BUFFERS ON the same query goes from ~ 20% to ~ 10% on top of actual runtime. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com> Discussion: https://www.postgresql.org/message-id/flat/CAP53PkxFP7i7-wy98ZmEJ11edYq-RrPvJoa4kzGhBBjERA4Nyw%40mail.gmail.com#e8dfd018a07d7f8d41565a079d40c564 fix up execprocnode 2
pg_stat_database's blk_read_time and blk_write_time were accumulated in the backend-local counters pgStatBlockReadTime and pgStatBlockWriteTime, which pgstat_count_io_op_time() incremented alongside the equivalent fields of the current instrumentation stack entry. With the stack-based instrumentation the session-level totals are available in instr_top, so the separate counters are redundant. Remove the counters and the pgstat_count_buffer_read_time() and pgstat_count_buffer_write_time() macros, and instead have pgstat_update_dbstats() compute the increase of the shared and local block read/write times in instr_top since the previous report. This covers the same activity as before: relation and temp relation reads, writes and extends, but not temporary file I/O or WAL I/O. instr_top is only updated when an instrumentation stack entry is finalized, but pgstat_report_stat() is only called when a backend is idle outside of a transaction block, after each autovacuum relation, or at process exit, all points where the stack is empty, so nothing is missed. Parallel workers account for their own activity in their own instr_top and report it themselves, like before. The leader additionally imports the workers' usage into its stack for query-level consumers such as EXPLAIN and pg_stat_statements, and thus into its instr_top. To count each activity exactly once, track everything imported from workers in the new instr_from_workers, updated at the same points as the import, and subtract it from instr_top when reporting. Since the amounts are now computed as a difference of the session totals, these totals must never be reset. Author: Lukas Fittl <lukas@fittl.com> Reviewed-by: Discussion:
This sets up a separate instrumentation stack that is used whilst an Index Scan or Index Only Scan does scanning on the table, for example due to additional data being needed. EXPLAIN ANALYZE will now show "Table Buffers" that represent such activity. The activity is also included in regular "Buffers" together with index activity and that of any child nodes. Author: Lukas Fittl <lukas@fittl.com> Suggested-by: Andres Freund <andres@anarazel.de> Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com> Reviewed-by: Tomas Vondra <tomas@vondra.me> Discussion: https://www.postgresql.org/message-id/flat/CAP53PkxrmpECzVFpeeEEHDGe6u625s%2BYkmVv5-gw3L_NDSfbiA%40mail.gmail.com#cb583a08e8e096aa1f093bb178906173
This is intended for testing instrumentation related logic as it pertains to the top level stack that is maintained as a running total. There is currently no in-core user that utilizes the top-level values in this manner, and especially during abort situations this helps ensure we don't lose activity due to incorrect handling of unfinalized node stacks.
lfittl
force-pushed
the
instrument-usage-stack-v18
branch
from
October 3, 2026 17:01
cbcf577 to
8ef30a4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.