REPACK (CONCURRENTLY) can't complete after ~105M concurrent updates/deletes
Hackorum builds and tests every patch posted to the lists, not only commitfest submissions. This is Hackorum's own CI rather than the PostgreSQL project's, and it is still under testing - please report anything that looks wrong.
You can run a PostgreSQL built from this patch straight from Docker, with no checkout and no build:
docker run --rm -p 5432:5432 ghcr.io/hackorum-dev/postgres-patch:t253948psql -h localhost -U postgresBuilt from patchset v6 (message #6), October 06, 2026 at 08:02 AM.
Every patchset is also pushed to a branch of our PostgreSQL fork, so you can check out the same tree CI built. Without a PostgreSQL checkout:
git clone --branch t253948_6 https://github.com/hackorum-dev/postgres.gitIn a checkout you already have, add the fork once:
git remote add hackorum https://github.com/hackorum-dev/postgres.gitthen, for this patchset and every later one:
git fetch hackorum t253948_6 && git checkout t253948_6Patchset v6 (message #6) is on t253948_6
Hey,
specifically CC'ing Antonin as we already talked about some
squeezing/repacking problems in past.
As it's one of the features I'm looking forward to most, I did quite a lot
of testing of
REPACK (CONCURRENTLY) over the last 2 weeks. I'm happy to say I wasn't able
to hit any show stopper, no matter how much I tried to break it (although I
have some edge case scenarios for later). I ran 75+ hostile runs (lots of
them scenarios I've previously seen hurt pg_repack/pg_squeeze).
The one thing I found is a hard limit during the catch-up. I noticed the
backend memory growing with the amount of concurrent updates, and when I
pushed it with a big batch update running concurrently, REPACK
(CONCURRENTLY) got OOM-killed, and with enough memory it failed on a hard
limit instead.
As far as I can tell this came over from pg_squeeze, which applies changes
the same way. pg_repack doesn't hit this particular scenario.
To verify it I tried multiple outcomes and got to
limit outcome no of changes
1 GB backend SIGKILLed & cluster crash restart ~18.0M
4 GB backend SIGKILLed & cluster crash restart ~84.3M
8 GB ERROR: invalid memory alloc request 104,820,740
size 1677721600
To make it deterministic I paused REPACK just before catch-up (1M-row
table) and ran N full-table updates from another session, then let it go.
On a table that small REPACK would otherwise finish long before enough
changes pile up; on a big table the copy and index builds take hours, which
gives the same effect without any trick.
Memory limit is --memory / --memory-swap set the same on a docker
container, release build. All three on master on Apple Silicon; the 1 GB
case I repeated on 19beta4, both on Apple Silicon and in the same container
on amd64 VM (~17.2M and ~16.8M), so it's not master or ARM specific.
Memory increase is linear, somewhere around 50 bytes per replayed
update/delete, and no GUC caps it. Only when I treid 8 GB run I was
surprised by the fact it hits the fixed number of changes.
The surprise is that REPACK (CONCURRENTLY) can't finish with more than 105M
rows updated/deleted concurrently (rows, not statements). Imagine something
with large number of HOT updates and other adverse condition. It's not
going to affect basic use cases, but if I think about tables where I would
see REPACK (CONCURRENTLY) used as alternative to non-blocking CLUSTER
during the I/O problems due to the data collocation, this is actually quite
realistic. Over last year alone there was more than handful scenarios where
this was unfortunately peak time solution to get data sorted. While it
might be considered abuse, imo it's legitimate.
The example would be REPACK of 250 - 500 GB table (don't even get me
started on
over-indexed ones). Just this week I dealt with a 100 GB table where full
operation on managed instance would take roughly 1.5 hours, that's already
only ~19k row changes/s. Get to 5h REPACK and all it takes is ~6k row
changes/s. Not every day problem, but you know how it goes - when it
rains...
DISCLAIMER: what follows was LLM assisted. The numbers line up exactly with
the ERROR I got, but I can't claim I came up with the explanation myself.
---
Why it happens: every tuple in the new heap is inserted by the REPACK
transaction, and apply_concurrent_changes() does a CommandCounterIncrement
before each replayed UPDATE or DELETE. So each of those modifies a tuple
with our own xmin and an older cmin, which means a new combo CID per
change, kept until commit. It shows up as growth in "Combo CIDs" in
pg_log_backend_memory_contexts().
Where the ceiling comes from: combocid.c starts the array at 100 entries
and doubles it, and after 100 * 2^20 = 104,857,600 entries the next
repalloc (1,677,721,600 bytes) exceeds MaxAllocSize. That's independent
of available memory, so the limit is the same everywhere.
---
I believe this is not a stopper for REPACK (CONCURRENTLY) but given the
visibility of the feature this might be thing that migth get documented. It
will also attrack people who might not have prior experience with
concurrent repacking tools. Hence we can only hope the REPACKing is done in
sane periods, but then as written above - I definitely used pg_squeeze in
past to solve data layout issues. At the same time this might scale up to 5
GB more memory needed in times when DBAs might be already facing adverse
conditions.
Hopefully over the wekeend I'm going to publish the findings on my site (
boringsql.com) to document this behaviour under title "Sizing REPACK
(CONCURRENTLY) for busy tables".
Hope this make sense
Radim
PS: during my runs I also replicated the issue Thom Brown reported with
TOAST table.
Hi Radim,
Your analysis is right. On master, 2M replayed UPDATEs used about 106MB
for combo CIDs, close to your number.
Patch 0008 in Antonin's "REPACK enhancements" series [1]https://www.google.com/url?q=https://postgr.es/m/224072.1789577949@localhost&source=gmail&ust=1790462338814000&sa=E removes the
problem. It replays changes with their original XIDs and without the
per-change CommandCounterIncrement(). The same test made no combo CIDs
there. That series is aimed at 20, but I am also not sure if it is
appropriate to
say repack have performance issue in document.
I used an LLM for this as well. The numbers are from a run on my machine.
[1]: https://www.google.com/url?q=https://postgr.es/m/224072.1789577949@localhost&source=gmail&ust=1790462338814000&sa=E
https://www.google.com/url?q=https://postgr.es/m/224072.1789577949@localhost&source=gmail&ust=1790462338814000&sa=E
Thanks
Shihao
On Fri, Sep 25, 2026 03:59 PM, Radim Marek <radim@boringsql.com> wrote:
Show quoted text
Hey,
specifically CC'ing Antonin as we already talked about some
squeezing/repacking problems in past.As it's one of the features I'm looking forward to most, I did quite a lot
of testing of
REPACK (CONCURRENTLY) over the last 2 weeks. I'm happy to say I wasn't
able to hit any show stopper, no matter how much I tried to break it
(although I have some edge case scenarios for later). I ran 75+ hostile
runs (lots of them scenarios I've previously seen hurt
pg_repack/pg_squeeze).The one thing I found is a hard limit during the catch-up. I noticed the
backend memory growing with the amount of concurrent updates, and when I
pushed it with a big batch update running concurrently, REPACK
(CONCURRENTLY) got OOM-killed, and with enough memory it failed on a hard
limit instead.As far as I can tell this came over from pg_squeeze, which applies changes
the same way. pg_repack doesn't hit this particular scenario.To verify it I tried multiple outcomes and got to
limit outcome no of changes
1 GB backend SIGKILLed & cluster crash restart ~18.0M
4 GB backend SIGKILLed & cluster crash restart ~84.3M
8 GB ERROR: invalid memory alloc request 104,820,740
size 1677721600To make it deterministic I paused REPACK just before catch-up (1M-row
table) and ran N full-table updates from another session, then let it go.
On a table that small REPACK would otherwise finish long before enough
changes pile up; on a big table the copy and index builds take hours, which
gives the same effect without any trick.Memory limit is --memory / --memory-swap set the same on a docker
container, release build. All three on master on Apple Silicon; the 1 GB
case I repeated on 19beta4, both on Apple Silicon and in the same container
on amd64 VM (~17.2M and ~16.8M), so it's not master or ARM specific.Memory increase is linear, somewhere around 50 bytes per replayed
update/delete, and no GUC caps it. Only when I treid 8 GB run I was
surprised by the fact it hits the fixed number of changes.The surprise is that REPACK (CONCURRENTLY) can't finish with more than
105M rows updated/deleted concurrently (rows, not statements). Imagine
something with large number of HOT updates and other adverse condition.
It's not going to affect basic use cases, but if I think about tables where
I would see REPACK (CONCURRENTLY) used as alternative to non-blocking
CLUSTER during the I/O problems due to the data collocation, this is
actually quite realistic. Over last year alone there was more than handful
scenarios where this was unfortunately peak time solution to get data
sorted. While it might be considered abuse, imo it's legitimate.The example would be REPACK of 250 - 500 GB table (don't even get me
started on
over-indexed ones). Just this week I dealt with a 100 GB table where full
operation on managed instance would take roughly 1.5 hours, that's already
only ~19k row changes/s. Get to 5h REPACK and all it takes is ~6k row
changes/s. Not every day problem, but you know how it goes - when it
rains...DISCLAIMER: what follows was LLM assisted. The numbers line up exactly
with the ERROR I got, but I can't claim I came up with the explanation
myself.---
Why it happens: every tuple in the new heap is inserted by the REPACK
transaction, and apply_concurrent_changes() does a CommandCounterIncrement
before each replayed UPDATE or DELETE. So each of those modifies a tuple
with our own xmin and an older cmin, which means a new combo CID per
change, kept until commit. It shows up as growth in "Combo CIDs" in
pg_log_backend_memory_contexts().Where the ceiling comes from: combocid.c starts the array at 100 entries
and doubles it, and after 100 * 2^20 = 104,857,600 entries the next
repalloc (1,677,721,600 bytes) exceeds MaxAllocSize. That's independent
of available memory, so the limit is the same everywhere.---
I believe this is not a stopper for REPACK (CONCURRENTLY) but given the
visibility of the feature this might be thing that migth get documented. It
will also attrack people who might not have prior experience with
concurrent repacking tools. Hence we can only hope the REPACKing is done in
sane periods, but then as written above - I definitely used pg_squeeze in
past to solve data layout issues. At the same time this might scale up to 5
GB more memory needed in times when DBAs might be already facing adverse
conditions.Hopefully over the wekeend I'm going to publish the findings on my site (
boringsql.com) to document this behaviour under title "Sizing REPACK
(CONCURRENTLY) for busy tables".Hope this make sense
Radim
PS: during my runs I also replicated the issue Thom Brown reported with
TOAST table.
Hi Shihao, thank you for the confirmation. I wasn't aware of the
enhancements series.
It's definitely not a performance issue in the doc; rathe a resource limit.
Since 19 will go with this limitation I attached a small doc patch about
this. Hope I got the use of 'other' correctly based on current version
https://www.postgresql.org/docs/19/sql-repack.html
Radim
On Sat, 26 Sept 2026 at 00:49, shihao zhong <zhong950419@gmail.com> wrote:
Show quoted text
Hi Radim,
Your analysis is right. On master, 2M replayed UPDATEs used about 106MB
for combo CIDs, close to your number.Patch 0008 in Antonin's "REPACK enhancements" series [1] removes the
problem. It replays changes with their original XIDs and without the
per-change CommandCounterIncrement(). The same test made no combo CIDs
there. That series is aimed at 20, but I am also not sure if it is
appropriate to
say repack have performance issue in document.I used an LLM for this as well. The numbers are from a run on my machine.
Thanks
ShihaoOn Fri, Sep 25, 2026 03:59 PM, Radim Marek <radim@boringsql.com> wrote:
Hey,
specifically CC'ing Antonin as we already talked about some
squeezing/repacking problems in past.As it's one of the features I'm looking forward to most, I did quite a
lot of testing of
REPACK (CONCURRENTLY) over the last 2 weeks. I'm happy to say I wasn't
able to hit any show stopper, no matter how much I tried to break it
(although I have some edge case scenarios for later). I ran 75+ hostile
runs (lots of them scenarios I've previously seen hurt
pg_repack/pg_squeeze).The one thing I found is a hard limit during the catch-up. I noticed the
backend memory growing with the amount of concurrent updates, and when I
pushed it with a big batch update running concurrently, REPACK
(CONCURRENTLY) got OOM-killed, and with enough memory it failed on a hard
limit instead.As far as I can tell this came over from pg_squeeze, which applies
changes the same way. pg_repack doesn't hit this particular scenario.To verify it I tried multiple outcomes and got to
limit outcome no of changes
1 GB backend SIGKILLed & cluster crash restart ~18.0M
4 GB backend SIGKILLed & cluster crash restart ~84.3M
8 GB ERROR: invalid memory alloc request 104,820,740
size 1677721600To make it deterministic I paused REPACK just before catch-up (1M-row
table) and ran N full-table updates from another session, then let it go.
On a table that small REPACK would otherwise finish long before enough
changes pile up; on a big table the copy and index builds take hours, which
gives the same effect without any trick.Memory limit is --memory / --memory-swap set the same on a docker
container, release build. All three on master on Apple Silicon; the 1 GB
case I repeated on 19beta4, both on Apple Silicon and in the same container
on amd64 VM (~17.2M and ~16.8M), so it's not master or ARM specific.Memory increase is linear, somewhere around 50 bytes per replayed
update/delete, and no GUC caps it. Only when I treid 8 GB run I was
surprised by the fact it hits the fixed number of changes.The surprise is that REPACK (CONCURRENTLY) can't finish with more than
105M rows updated/deleted concurrently (rows, not statements). Imagine
something with large number of HOT updates and other adverse condition.
It's not going to affect basic use cases, but if I think about tables where
I would see REPACK (CONCURRENTLY) used as alternative to non-blocking
CLUSTER during the I/O problems due to the data collocation, this is
actually quite realistic. Over last year alone there was more than handful
scenarios where this was unfortunately peak time solution to get data
sorted. While it might be considered abuse, imo it's legitimate.The example would be REPACK of 250 - 500 GB table (don't even get me
started on
over-indexed ones). Just this week I dealt with a 100 GB table where full
operation on managed instance would take roughly 1.5 hours, that's already
only ~19k row changes/s. Get to 5h REPACK and all it takes is ~6k row
changes/s. Not every day problem, but you know how it goes - when it
rains...DISCLAIMER: what follows was LLM assisted. The numbers line up exactly
with the ERROR I got, but I can't claim I came up with the explanation
myself.---
Why it happens: every tuple in the new heap is inserted by the REPACK
transaction, and apply_concurrent_changes() does a CommandCounterIncrement
before each replayed UPDATE or DELETE. So each of those modifies a tuple
with our own xmin and an older cmin, which means a new combo CID per
change, kept until commit. It shows up as growth in "Combo CIDs" in
pg_log_backend_memory_contexts().Where the ceiling comes from: combocid.c starts the array at 100 entries
and doubles it, and after 100 * 2^20 = 104,857,600 entries the next
repalloc (1,677,721,600 bytes) exceeds MaxAllocSize. That's independent
of available memory, so the limit is the same everywhere.---
I believe this is not a stopper for REPACK (CONCURRENTLY) but given the
visibility of the feature this might be thing that migth get documented. It
will also attrack people who might not have prior experience with
concurrent repacking tools. Hence we can only hope the REPACKing is done in
sane periods, but then as written above - I definitely used pg_squeeze in
past to solve data layout issues. At the same time this might scale up to 5
GB more memory needed in times when DBAs might be already facing adverse
conditions.Hopefully over the wekeend I'm going to publish the findings on my site (
boringsql.com) to document this behaviour under title "Sizing REPACK
(CONCURRENTLY) for busy tables".Hope this make sense
Radim
PS: during my runs I also replicated the issue Thom Brown reported with
TOAST table.
Radim Marek <radim@boringsql.com> wrote:
Hi Shihao, thank you for the confirmation. I wasn't aware of the enhancements series.
It's definitely not a performance issue in the doc; rathe a resource limit. Since 19 will go with this limitation I attached a small doc patch
about this. Hope I got the use of 'other' correctly based on current version https://www.postgresql.org/docs/19/sql-repack.html
It's unfortunate that REPACK is probably the only command that exercises this
combo CID limit. However, there can be other limits, e.g. REPACK being unable
to catch up if the change rate is too high, excessive usage of disk for the
decoded changes, etc.
I think that the most important thing for the user to know is that the purpose
of REPACK is to improve performance (via bloat removal and/or clustering),
whereas "failsafe VACUUM" (i.e. VACUUM with INDEX_CLEANUP set to off) should
be used to avoid the risk of XID wraparound. This kind of VACCUM is probably
much faster and less likely to fail.
I don't have a good idea about the wording right now.
--
Antonin Houska
Web: https://www.cybertec-postgresql.com
Hi Antonin,
I think that the most important thing for the user to know is that the
purpose
of REPACK is to improve performance (via bloat removal and/or clustering),
whereas "failsafe VACUUM" (i.e. VACUUM with INDEX_CLEANUP set to off)
should
be used to avoid the risk of XID wraparound.
I don't have a good idea about the wording right now.
Here is a try with Fable, as v2 of Radim's patch.
It adds a paragraph to Notes. REPACK is for bloat and clustering. For
wraparound uses VACUUM, because REPACK takes much longer and can fail
late.
Thanks,
Shihao
On 2026-Sep-27, shihao zhong wrote:
Here is a try with Fable, as v2 of Radim's patch.
It adds a paragraph to Notes. REPACK is for bloat and clustering. For
wraparound uses VACUUM, because REPACK takes much longer and can fail
late.
Yeah, that sounds appropriate.
I think the original <note> paragraph is worth rewriting more deeply
though rather than just adding one more paragraph; I think some of the
things it mentions are not so relevant from the user's POV, and also I
think we can use some small changes elsewhere in the page.
What do you think of the attached? I used -U7 in `git format-patch` so
that the surrounding can be read directly from the patch.
--
Álvaro Herrera 48°01'N 7°57'E — https://www.EnterpriseDB.com/
Álvaro Herrera <alvherre@kurilemu.de> wrote:
On 2026-Sep-27, shihao zhong wrote:
Here is a try with Fable, as v2 of Radim's patch.
It adds a paragraph to Notes. REPACK is for bloat and clustering. For
wraparound uses VACUUM, because REPACK takes much longer and can fail
late.Yeah, that sounds appropriate.
I think the original <note> paragraph is worth rewriting more deeply
though rather than just adding one more paragraph; I think some of the
things it mentions are not so relevant from the user's POV, and also I
think we can use some small changes elsewhere in the page.What do you think of the attached? I used -U7 in `git format-patch` so
that the surrounding can be read directly from the patch.
I thought about this documentation update last week and wasn't sure it needs
to be that exact about how much memory is consumed per changed tuple. I think
we should rather encourage users to do monitoring in genaral: besides memory,
the system can run out of disk space (REPACK w/o CONCURRENTLY also creates a
copy of the table, but it's probably not used for big tables due to the
stronger lock). Besides that, REPACK runs in a single transaction, so it might
hold the xmin horizon(s) for too long (like any other long-running
transactions). Finally, the processing of the concurrent changes might not be
fast enough, in which case the AccessExclusiveLock may be needed for
surprisingly long time.
The paragraph about REPACK vs VACUUM LGTM.
--
Antonin Houska
Web: https://www.cybertec-postgresql.com