* [PATCHSET 0/2] Fix memlock account for kernel backed memory
@ 2026-10-07 21:38 Jens Axboe
2026-10-07 21:38 ` [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses" Jens Axboe
2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
0 siblings, 2 replies; 3+ messages in thread
From: Jens Axboe @ 2026-10-07 21:38 UTC (permalink / raw)
To: io-uring; +Cc: dw, hengyul
Hi,
Traditionally, kernel allocated memory did not account against memlock
limits, but a change in 6.14 inadvertently broke that. Reinstate the old
behavior.
Split into 2 patches for ease of backporting, and based on this original
patch and report:
https://lore.kernel.org/io-uring/20261006125732.3425762-1-hengyul@cs.unc.edu/
--
Jens Axboe
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses"
2026-10-07 21:38 [PATCHSET 0/2] Fix memlock account for kernel backed memory Jens Axboe
@ 2026-10-07 21:38 ` Jens Axboe
2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
1 sibling, 0 replies; 3+ messages in thread
From: Jens Axboe @ 2026-10-07 21:38 UTC (permalink / raw)
To: io-uring; +Cc: dw, hengyul, Jens Axboe
This reverts commit f12f0234cc14886bcfd53ffb7c8df4216dad51f1.
The next patch stops charging kernel allocated regions to RLIMIT_MEMLOCK
altogether, which leaves nothing for this to account. Revert it first so
that fix applies cleanly to the stable trees that need it, none of which
have this commit.
Signed-off-by: Jens Axboe <axboe@kernel.dk>
---
io_uring/memmap.c | 37 +++++++------------------------------
1 file changed, 7 insertions(+), 30 deletions(-)
diff --git a/io_uring/memmap.c b/io_uring/memmap.c
index 48c0eb012412..23e8a85111bc 100644
--- a/io_uring/memmap.c
+++ b/io_uring/memmap.c
@@ -16,10 +16,8 @@
#include "zcrx.h"
static bool io_mem_alloc_compound(struct page **pages, int nr_pages,
- size_t size, gfp_t gfp,
- struct user_struct *user)
+ size_t size, gfp_t gfp)
{
- unsigned long nr_compound, extra;
struct page *page;
int i, order;
@@ -29,22 +27,9 @@ static bool io_mem_alloc_compound(struct page **pages, int nr_pages,
else if (order)
gfp |= __GFP_COMP;
- /*
- * get_order() rounds a non power of two size up, so the allocation
- * can hold more pages than the region exposes. Account those too,
- * and leave the compound allocation alone if they do not fit.
- */
- nr_compound = 1UL << order;
- extra = nr_compound - nr_pages;
- if (extra && user && __io_account_mem(user, extra))
- return false;
-
page = alloc_pages(gfp, order);
- if (!page) {
- if (extra && user)
- __io_unaccount_mem(user, extra);
+ if (!page)
return false;
- }
for (i = 0; i < nr_pages; i++)
pages[i] = page + i;
@@ -120,15 +105,8 @@ void io_free_region(struct user_struct *user, struct io_mapped_region *mr)
}
if ((mr->flags & IO_REGION_F_VMAP) && mr->ptr)
vunmap(mr->ptr);
- if (mr->nr_pages && user) {
- unsigned long nr_accounted = mr->nr_pages;
-
- /* a compound region was accounted for the whole allocation */
- if (mr->flags & IO_REGION_F_SINGLE_REF)
- nr_accounted = 1UL << get_order(io_region_size(mr));
-
- __io_unaccount_mem(user, nr_accounted);
- }
+ if (mr->nr_pages && user)
+ __io_unaccount_mem(user, mr->nr_pages);
memset(mr, 0, sizeof(*mr));
}
@@ -173,8 +151,7 @@ static int io_region_pin_pages(struct io_mapped_region *mr,
static int io_region_allocate_pages(struct io_mapped_region *mr,
struct io_uring_region_desc *reg,
- unsigned long mmap_offset,
- struct user_struct *user)
+ unsigned long mmap_offset)
{
gfp_t gfp = GFP_KERNEL_ACCOUNT | __GFP_ZERO | __GFP_NOWARN;
size_t size = io_region_size(mr);
@@ -185,7 +162,7 @@ static int io_region_allocate_pages(struct io_mapped_region *mr,
if (!pages)
return -ENOMEM;
- if (io_mem_alloc_compound(pages, mr->nr_pages, size, gfp, user)) {
+ if (io_mem_alloc_compound(pages, mr->nr_pages, size, gfp)) {
mr->flags |= IO_REGION_F_SINGLE_REF;
goto done;
}
@@ -240,7 +217,7 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
if (reg->flags & IORING_MEM_REGION_TYPE_USER)
ret = io_region_pin_pages(mr, reg);
else
- ret = io_region_allocate_pages(mr, reg, mmap_offset, ctx->user);
+ ret = io_region_allocate_pages(mr, reg, mmap_offset);
if (ret)
goto out_free;
--
2.55.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
* [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK
2026-10-07 21:38 [PATCHSET 0/2] Fix memlock account for kernel backed memory Jens Axboe
2026-10-07 21:38 ` [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses" Jens Axboe
@ 2026-10-07 21:38 ` Jens Axboe
1 sibling, 0 replies; 3+ messages in thread
From: Jens Axboe @ 2026-10-07 21:38 UTC (permalink / raw)
To: io-uring; +Cc: dw, hengyul, Jens Axboe, stable
io_create_region() charges every region to RLIMIT_MEMLOCK. This was
correct when the only region type was the user provided parameter region, but the SQ/CQ
rings and provided buffer rings have since been converted to regions as
well. Those we have tradionally excluded from that. Kernel allocated
ring memory is memcg accounted, and never counted against the memlock
limit. Since 6.14 an unprivileged user runs out of the 8MB default after
a couple of dozen rings.
Charge RLIMIT_MEMLOCK only for regions backed by pinned user memory, which
is what the limit is for, and leave kernel allocations to memcg. User
provided ring memory keeps being charged, it is pinned.
Reported-by: Hengyu Liang <hengyul@cs.unc.edu>
Link: https://lore.kernel.org/io-uring/20261006125732.3425762-1-hengyul@cs.unc.edu/
Fixes: 8078486e1d53 ("io_uring: use region api for SQ")
Fixes: 81a4058e0cd0 ("io_uring: use region api for CQ")
Fixes: ef62de3c4ad5 ("io_uring/kbuf: use region api for pbuf rings")
Cc: stable@vger.kernel.org
Signed-off-by: Jens Axboe <axboe@kernel.dk>
---
io_uring/memmap.c | 28 +++++++++++++++++++---------
1 file changed, 19 insertions(+), 9 deletions(-)
diff --git a/io_uring/memmap.c b/io_uring/memmap.c
index 23e8a85111bc..da2328b52b38 100644
--- a/io_uring/memmap.c
+++ b/io_uring/memmap.c
@@ -105,7 +105,7 @@ void io_free_region(struct user_struct *user, struct io_mapped_region *mr)
}
if ((mr->flags & IO_REGION_F_VMAP) && mr->ptr)
vunmap(mr->ptr);
- if (mr->nr_pages && user)
+ if ((mr->flags & IO_REGION_F_USER_PROVIDED) && user)
__io_unaccount_mem(user, mr->nr_pages);
memset(mr, 0, sizeof(*mr));
@@ -132,11 +132,12 @@ static int io_region_init_ptr(struct io_mapped_region *mr)
}
static int io_region_pin_pages(struct io_mapped_region *mr,
- struct io_uring_region_desc *reg)
+ struct io_uring_region_desc *reg,
+ struct user_struct *user)
{
size_t size = io_region_size(mr);
struct page **pages;
- int nr_pages;
+ int nr_pages, ret;
pages = io_pin_pages(reg->user_addr, size, &nr_pages);
if (IS_ERR(pages))
@@ -144,6 +145,16 @@ static int io_region_pin_pages(struct io_mapped_region *mr,
if (WARN_ON_ONCE(nr_pages != mr->nr_pages))
return -EFAULT;
+ /* pinned user memory is what RLIMIT_MEMLOCK is for */
+ if (user) {
+ ret = __io_account_mem(user, nr_pages);
+ if (ret) {
+ unpin_user_pages(pages, nr_pages);
+ kvfree(pages);
+ return ret;
+ }
+ }
+
mr->pages = pages;
mr->flags |= IO_REGION_F_USER_PROVIDED;
return 0;
@@ -207,15 +218,14 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
return -EOVERFLOW;
nr_pages = reg->size >> PAGE_SHIFT;
- if (ctx->user) {
- ret = __io_account_mem(ctx->user, nr_pages);
- if (ret)
- return ret;
- }
mr->nr_pages = nr_pages;
+ /*
+ * Only pinned user memory counts against RLIMIT_MEMLOCK, kernel
+ * allocated regions are memcg accounted through GFP_KERNEL_ACCOUNT.
+ */
if (reg->flags & IORING_MEM_REGION_TYPE_USER)
- ret = io_region_pin_pages(mr, reg);
+ ret = io_region_pin_pages(mr, reg, ctx->user);
else
ret = io_region_allocate_pages(mr, reg, mmap_offset);
if (ret)
--
2.55.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-07 21:39 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 21:38 [PATCHSET 0/2] Fix memlock account for kernel backed memory Jens Axboe
2026-10-07 21:38 ` [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses" Jens Axboe
2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox