We run Pgpool-II in streaming replication mode in front of PostgreSQL 14, one
primary and one standby. On one of our pools, child processes were dying by
SIGSEGV continuously, a few per minute, on 4.4.13, 4.5.8 and 4.7.2 alike. The
client on that pool uses named prepared statements (Go pgx with its statement
cache). The relevant settings are load_balance_mode = on,
statement_level_load_balance = on and
disable_load_balance_on_write = transaction. We captured core dumps from an
unstripped 4.7.2 build.
#0 is_select_query (node=0x7da878105980,
sql=0x7da878105af0 <error: Cannot access memory at address 0x7da878105af0>)
at protocol/pool_process_query.c:1156
#1 pool_pending_message_dest_set (message=0x7da878103d78,
query_context=0x7da878193c58) at context/pool_session_context.c:1134
#2 Bind (frontend=0x7da88d452b48, backend=0x7da88d449248,
contents=0x7da88d43d438 "") at protocol/pool_proto_modules.c:1900
#3 ProcessFrontendResponse (...) at protocol/pool_proto_modules.c:2974
#4 read_packets_and_process (...) at protocol/pool_process_query.c:5142
#5 pool_process_query (...) at protocol/pool_process_query.c:299
#6 do_child (fds=0x7da88d45ce30) at protocol/child.c:464
The query context itself is intact, its MemoryContext header still reads
QueryContextMemoryContext, but its original_query and parse_tree point
into an unmapped page. A second core crashed in function_call_walker() via
pool_has_function_call() from Execute(), with the parse tree node pointer
reading 0x656664666631352d, which is ASCII -51ffdfe: a fragment of a bind
parameter that had been written over the freed tree.
parse_before_bind() handles a Bind for a named statement that was parsed on a
load balance node and now has to run on the primary. It builds the modified
query context with pool_query_context_shallow_copy(), which memcpy()s the
source and therefore shares the source's original_query and parse_tree
pointers, and then repoints the Parse and Bind sent messages at the copy. After
that the source is referenced by nothing except the previous unnamed portal.
Back in Bind(), pool_add_sent_message() for the new portal removes that old
portal, can_query_context_destroy() finds the source referenced once, and the
source is destroyed, freeing the parse tree and query string the copy still
points at. The next Bind or Execute of that named statement walks freed memory.
pool_query_context_shallow_copy() is unchanged on master.
Copying original_query, rewritten_query, parse_tree,
rewritten_parse_tree and query_w_hex into the new context's own memory
context fixes it. With that change, an environment that had been logging child
segfaults in nearly every minute has run clean, and the backend's matching
unexpected EOF on client connection with an open transaction messages went to
zero at the same moment. I will attach the patch, which applies to 4.3 through
master, as a comment.
We run Pgpool-II in streaming replication mode in front of PostgreSQL 14, one
primary and one standby. On one of our pools, child processes were dying by
SIGSEGV continuously, a few per minute, on 4.4.13, 4.5.8 and 4.7.2 alike. The
client on that pool uses named prepared statements (Go pgx with its statement
cache). The relevant settings are
load_balance_mode = on,statement_level_load_balance = onanddisable_load_balance_on_write = transaction. We captured core dumps from anunstripped 4.7.2 build.
The query context itself is intact, its MemoryContext header still reads
QueryContextMemoryContext, but itsoriginal_queryandparse_treepointinto an unmapped page. A second core crashed in
function_call_walker()viapool_has_function_call()fromExecute(), with the parse tree node pointerreading
0x656664666631352d, which is ASCII-51ffdfe: a fragment of a bindparameter that had been written over the freed tree.
parse_before_bind()handles a Bind for a named statement that was parsed on aload balance node and now has to run on the primary. It builds the modified
query context with
pool_query_context_shallow_copy(), whichmemcpy()s thesource and therefore shares the source's
original_queryandparse_treepointers, and then repoints the Parse and Bind sent messages at the copy. After
that the source is referenced by nothing except the previous unnamed portal.
Back in
Bind(),pool_add_sent_message()for the new portal removes that oldportal,
can_query_context_destroy()finds the source referenced once, and thesource is destroyed, freeing the parse tree and query string the copy still
points at. The next Bind or Execute of that named statement walks freed memory.
pool_query_context_shallow_copy()is unchanged on master.Copying
original_query,rewritten_query,parse_tree,rewritten_parse_treeandquery_w_hexinto the new context's own memorycontext fixes it. With that change, an environment that had been logging child
segfaults in nearly every minute has run clean, and the backend's matching
unexpected EOF on client connection with an open transactionmessages went tozero at the same moment. I will attach the patch, which applies to 4.3 through
master, as a comment.