public inbox for io-uring@vger.kernel.org
 help / color / mirror / Atom feed
* [PATCHSET 0/2] Fix memlock account for kernel backed memory
@ 2026-10-07 21:38 Jens Axboe
  2026-10-07 21:38 ` [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses" Jens Axboe
  2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
  0 siblings, 2 replies; 7+ messages in thread
From: Jens Axboe @ 2026-10-07 21:38 UTC (permalink / raw)
  To: io-uring; +Cc: dw, hengyul

Hi,

Traditionally, kernel allocated memory did not account against memlock
limits, but a change in 6.14 inadvertently broke that. Reinstate the old
behavior.

Split into 2 patches for ease of backporting, and based on this original
patch and report:

https://lore.kernel.org/io-uring/20261006125732.3425762-1-hengyul@cs.unc.edu/

-- 
Jens Axboe


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses"
  2026-10-07 21:38 [PATCHSET 0/2] Fix memlock account for kernel backed memory Jens Axboe
@ 2026-10-07 21:38 ` Jens Axboe
  2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
  1 sibling, 0 replies; 7+ messages in thread
From: Jens Axboe @ 2026-10-07 21:38 UTC (permalink / raw)
  To: io-uring; +Cc: dw, hengyul, Jens Axboe

This reverts commit f12f0234cc14886bcfd53ffb7c8df4216dad51f1.

The next patch stops charging kernel allocated regions to RLIMIT_MEMLOCK
altogether, which leaves nothing for this to account. Revert it first so
that fix applies cleanly to the stable trees that need it, none of which
have this commit.

Signed-off-by: Jens Axboe <axboe@kernel.dk>
---
 io_uring/memmap.c | 37 +++++++------------------------------
 1 file changed, 7 insertions(+), 30 deletions(-)

diff --git a/io_uring/memmap.c b/io_uring/memmap.c
index 48c0eb012412..23e8a85111bc 100644
--- a/io_uring/memmap.c
+++ b/io_uring/memmap.c
@@ -16,10 +16,8 @@
 #include "zcrx.h"
 
 static bool io_mem_alloc_compound(struct page **pages, int nr_pages,
-				  size_t size, gfp_t gfp,
-				  struct user_struct *user)
+				  size_t size, gfp_t gfp)
 {
-	unsigned long nr_compound, extra;
 	struct page *page;
 	int i, order;
 
@@ -29,22 +27,9 @@ static bool io_mem_alloc_compound(struct page **pages, int nr_pages,
 	else if (order)
 		gfp |= __GFP_COMP;
 
-	/*
-	 * get_order() rounds a non power of two size up, so the allocation
-	 * can hold more pages than the region exposes. Account those too,
-	 * and leave the compound allocation alone if they do not fit.
-	 */
-	nr_compound = 1UL << order;
-	extra = nr_compound - nr_pages;
-	if (extra && user && __io_account_mem(user, extra))
-		return false;
-
 	page = alloc_pages(gfp, order);
-	if (!page) {
-		if (extra && user)
-			__io_unaccount_mem(user, extra);
+	if (!page)
 		return false;
-	}
 
 	for (i = 0; i < nr_pages; i++)
 		pages[i] = page + i;
@@ -120,15 +105,8 @@ void io_free_region(struct user_struct *user, struct io_mapped_region *mr)
 	}
 	if ((mr->flags & IO_REGION_F_VMAP) && mr->ptr)
 		vunmap(mr->ptr);
-	if (mr->nr_pages && user) {
-		unsigned long nr_accounted = mr->nr_pages;
-
-		/* a compound region was accounted for the whole allocation */
-		if (mr->flags & IO_REGION_F_SINGLE_REF)
-			nr_accounted = 1UL << get_order(io_region_size(mr));
-
-		__io_unaccount_mem(user, nr_accounted);
-	}
+	if (mr->nr_pages && user)
+		__io_unaccount_mem(user, mr->nr_pages);
 
 	memset(mr, 0, sizeof(*mr));
 }
@@ -173,8 +151,7 @@ static int io_region_pin_pages(struct io_mapped_region *mr,
 
 static int io_region_allocate_pages(struct io_mapped_region *mr,
 				    struct io_uring_region_desc *reg,
-				    unsigned long mmap_offset,
-				    struct user_struct *user)
+				    unsigned long mmap_offset)
 {
 	gfp_t gfp = GFP_KERNEL_ACCOUNT | __GFP_ZERO | __GFP_NOWARN;
 	size_t size = io_region_size(mr);
@@ -185,7 +162,7 @@ static int io_region_allocate_pages(struct io_mapped_region *mr,
 	if (!pages)
 		return -ENOMEM;
 
-	if (io_mem_alloc_compound(pages, mr->nr_pages, size, gfp, user)) {
+	if (io_mem_alloc_compound(pages, mr->nr_pages, size, gfp)) {
 		mr->flags |= IO_REGION_F_SINGLE_REF;
 		goto done;
 	}
@@ -240,7 +217,7 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
 	if (reg->flags & IORING_MEM_REGION_TYPE_USER)
 		ret = io_region_pin_pages(mr, reg);
 	else
-		ret = io_region_allocate_pages(mr, reg, mmap_offset, ctx->user);
+		ret = io_region_allocate_pages(mr, reg, mmap_offset);
 	if (ret)
 		goto out_free;
 
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK
  2026-10-07 21:38 [PATCHSET 0/2] Fix memlock account for kernel backed memory Jens Axboe
  2026-10-07 21:38 ` [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses" Jens Axboe
@ 2026-10-07 21:38 ` Jens Axboe
  2026-10-08  4:57   ` Hengyu Liang
  1 sibling, 1 reply; 7+ messages in thread
From: Jens Axboe @ 2026-10-07 21:38 UTC (permalink / raw)
  To: io-uring; +Cc: dw, hengyul, Jens Axboe, stable

io_create_region() charges every region to RLIMIT_MEMLOCK. This was
correct when the only region type was the user provided parameter region, but the SQ/CQ
rings and provided buffer rings have since been converted to regions as
well. Those we have tradionally excluded from that. Kernel allocated
ring memory is memcg accounted, and never counted against the memlock
limit. Since 6.14 an unprivileged user runs out of the 8MB default after
a couple of dozen rings.

Charge RLIMIT_MEMLOCK only for regions backed by pinned user memory, which
is what the limit is for, and leave kernel allocations to memcg. User
provided ring memory keeps being charged, it is pinned.

Reported-by: Hengyu Liang <hengyul@cs.unc.edu>
Link: https://lore.kernel.org/io-uring/20261006125732.3425762-1-hengyul@cs.unc.edu/
Fixes: 8078486e1d53 ("io_uring: use region api for SQ")
Fixes: 81a4058e0cd0 ("io_uring: use region api for CQ")
Fixes: ef62de3c4ad5 ("io_uring/kbuf: use region api for pbuf rings")
Cc: stable@vger.kernel.org
Signed-off-by: Jens Axboe <axboe@kernel.dk>
---
 io_uring/memmap.c | 28 +++++++++++++++++++---------
 1 file changed, 19 insertions(+), 9 deletions(-)

diff --git a/io_uring/memmap.c b/io_uring/memmap.c
index 23e8a85111bc..da2328b52b38 100644
--- a/io_uring/memmap.c
+++ b/io_uring/memmap.c
@@ -105,7 +105,7 @@ void io_free_region(struct user_struct *user, struct io_mapped_region *mr)
 	}
 	if ((mr->flags & IO_REGION_F_VMAP) && mr->ptr)
 		vunmap(mr->ptr);
-	if (mr->nr_pages && user)
+	if ((mr->flags & IO_REGION_F_USER_PROVIDED) && user)
 		__io_unaccount_mem(user, mr->nr_pages);
 
 	memset(mr, 0, sizeof(*mr));
@@ -132,11 +132,12 @@ static int io_region_init_ptr(struct io_mapped_region *mr)
 }
 
 static int io_region_pin_pages(struct io_mapped_region *mr,
-			       struct io_uring_region_desc *reg)
+			       struct io_uring_region_desc *reg,
+			       struct user_struct *user)
 {
 	size_t size = io_region_size(mr);
 	struct page **pages;
-	int nr_pages;
+	int nr_pages, ret;
 
 	pages = io_pin_pages(reg->user_addr, size, &nr_pages);
 	if (IS_ERR(pages))
@@ -144,6 +145,16 @@ static int io_region_pin_pages(struct io_mapped_region *mr,
 	if (WARN_ON_ONCE(nr_pages != mr->nr_pages))
 		return -EFAULT;
 
+	/* pinned user memory is what RLIMIT_MEMLOCK is for */
+	if (user) {
+		ret = __io_account_mem(user, nr_pages);
+		if (ret) {
+			unpin_user_pages(pages, nr_pages);
+			kvfree(pages);
+			return ret;
+		}
+	}
+
 	mr->pages = pages;
 	mr->flags |= IO_REGION_F_USER_PROVIDED;
 	return 0;
@@ -207,15 +218,14 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
 		return -EOVERFLOW;
 
 	nr_pages = reg->size >> PAGE_SHIFT;
-	if (ctx->user) {
-		ret = __io_account_mem(ctx->user, nr_pages);
-		if (ret)
-			return ret;
-	}
 	mr->nr_pages = nr_pages;
 
+	/*
+	 * Only pinned user memory counts against RLIMIT_MEMLOCK, kernel
+	 * allocated regions are memcg accounted through GFP_KERNEL_ACCOUNT.
+	 */
 	if (reg->flags & IORING_MEM_REGION_TYPE_USER)
-		ret = io_region_pin_pages(mr, reg);
+		ret = io_region_pin_pages(mr, reg, ctx->user);
 	else
 		ret = io_region_allocate_pages(mr, reg, mmap_offset);
 	if (ret)
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK
  2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
@ 2026-10-08  4:57   ` Hengyu Liang
  2026-10-08 14:54     ` Jens Axboe
  0 siblings, 1 reply; 7+ messages in thread
From: Hengyu Liang @ 2026-10-08  4:57 UTC (permalink / raw)
  To: axboe; +Cc: dw, hengyul, io-uring, stable

On 10/7/26 3:38 PM, Jens Axboe wrote:
> Charge RLIMIT_MEMLOCK only for regions backed by pinned user memory, which
> is what the limit is for, and leave kernel allocations to memcg. User
> provided ring memory keeps being charged, it is pinned.

Sorry for the late reply, and thanks for picking this up.

I tested both patches on top of v7.3-rc4. The test case from my patch
prints "64 rings" again, kernel allocated buffer rings are fixed as well,
the per-user locked_vm count is balanced after ring create, close and
resize, and the liburing tests give the same results as before, except
that read-before-exit.t passes again with the default limit.

Tested-by: Hengyu Liang <hengyul@cs.unc.edu>

One case is not restored. v6.13 did not charge rings created with
IORING_SETUP_NO_MMAP either. Number of rings created out of 64, as an
unprivileged user with the default 8 MiB limit, 4096 entries each:

                                     v6.13  v7.3-rc4  this series
  io_uring_queue_init()                 64        16           64
  io_uring_queue_init_mem()             64        21           21

PostgreSQL 18 is in the second row. It puts its rings into shared memory
with io_uring_queue_init_mem() whenever liburing has it [1], so I expect
the failure they reported to stay. I have not run PostgreSQL itself.

Would you take a patch on top that leaves the SQ/CQ rings uncharged for
user memory too, as in v6.13? I can send one, or test whatever you prefer.

Also, kernel allocated IORING_REGISTER_MEM_REGION regions are no longer
charged. They have been charged since they were added in v6.14. With the
series a user with an 8 MiB limit can register a 512 MiB one.

[1] https://github.com/postgres/postgres/blob/REL_18_STABLE/src/backend/storage/aio/method_io_uring.c

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK
  2026-10-08  4:57   ` Hengyu Liang
@ 2026-10-08 14:54     ` Jens Axboe
  2026-10-08 17:06       ` [PATCH] io_uring: do not charge user provided SQ/CQ rings " Hengyu Liang
  0 siblings, 1 reply; 7+ messages in thread
From: Jens Axboe @ 2026-10-08 14:54 UTC (permalink / raw)
  To: Hengyu Liang; +Cc: dw, io-uring, stable

On 10/7/26 10:57 PM, Hengyu Liang wrote:
> On 10/7/26 3:38 PM, Jens Axboe wrote:
>> Charge RLIMIT_MEMLOCK only for regions backed by pinned user memory, which
>> is what the limit is for, and leave kernel allocations to memcg. User
>> provided ring memory keeps being charged, it is pinned.
> 
> Sorry for the late reply, and thanks for picking this up.
> 
> I tested both patches on top of v7.3-rc4. The test case from my patch
> prints "64 rings" again, kernel allocated buffer rings are fixed as well,
> the per-user locked_vm count is balanced after ring create, close and
> resize, and the liburing tests give the same results as before, except
> that read-before-exit.t passes again with the default limit.
> 
> Tested-by: Hengyu Liang <hengyul@cs.unc.edu>

Thanks for testing!

> One case is not restored. v6.13 did not charge rings created with
> IORING_SETUP_NO_MMAP either. Number of rings created out of 64, as an
> unprivileged user with the default 8 MiB limit, 4096 entries each:
> 
>                                      v6.13  v7.3-rc4  this series
>   io_uring_queue_init()                 64        16           64
>   io_uring_queue_init_mem()             64        21           21
> 
> PostgreSQL 18 is in the second row. It puts its rings into shared memory
> with io_uring_queue_init_mem() whenever liburing has it [1], so I expect
> the failure they reported to stay. I have not run PostgreSQL itself.
> 
> Would you take a patch on top that leaves the SQ/CQ rings uncharged for
> user memory too, as in v6.13? I can send one, or test whatever you prefer.

Yes for sure, please feel free to send a patch we can apply on top.
>
> Also, kernel allocated IORING_REGISTER_MEM_REGION regions are no longer
> charged. They have been charged since they were added in v6.14. With the
> series a user with an 8 MiB limit can register a 512 MiB one.
> 
> [1] https://github.com/postgres/postgres/blob/REL_18_STABLE/src/backend/storage/aio/method_io_uring.c

We should still account IORING_MEM_REGION_TYPE_USER, anything that isn't
memcg accounted. I'll take a closer look, but if that isn't the case,
then yes that also needs a followup.

-- 
Jens Axboe

^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK
  2026-10-08 14:54     ` Jens Axboe
@ 2026-10-08 17:06       ` Hengyu Liang
  2026-10-08 19:11         ` Jens Axboe
  0 siblings, 1 reply; 7+ messages in thread
From: Hengyu Liang @ 2026-10-08 17:06 UTC (permalink / raw)
  To: Jens Axboe; +Cc: Pavel Begunkov, David Wei, io-uring, linux-kernel, netdev

Commit 8078486e1d53 ("io_uring: use region api for SQ") and commit
81a4058e0cd0 ("io_uring: use region api for CQ") made io_uring_setup()
create the rings with io_create_region().

However, io_create_region() charges user provided memory to
RLIMIT_MEMLOCK, and the rings of an IORING_SETUP_NO_MMAP ring were not
charged before those commits. As of now, a user without CAP_IPC_LOCK
gets ENOMEM from io_uring_queue_init_mem() when their rings exceed the
limit, which is 8 MiB by default. PostgreSQL 18 creates its rings with
this function [1].

The issue can be reproduced with a simple liburing program, run as an
unprivileged user:

    #include <liburing.h>
    #include <stdio.h>
    #include <sys/mman.h>

    int main(void)
    {
            static struct io_uring ring[64];
            char *mem = mmap(NULL, 64 << 20, PROT_READ | PROT_WRITE,
                             MAP_SHARED | MAP_ANONYMOUS, -1, 0);
            int i;

            for (i = 0; i < 64; i++) {
                    struct io_uring_params p = { };

                    if (io_uring_queue_init_mem(4096, &ring[i], &p,
                                                mem + (i << 20),
                                                1 << 20) < 0)
                            break;
            }
            printf("%d rings\n", i);
            return 0;
    }

Before those commits (v6.13), it prints "64 rings". After those commits
(v6.14), it prints "21 rings".

This patch makes io_create_region() take the user to charge, like
io_free_region() does, and passes no user for the SQ/CQ rings.

Link: https://github.com/postgres/postgres/blob/REL_18_STABLE/src/backend/storage/aio/method_io_uring.c [1]
Fixes: 8078486e1d53 ("io_uring: use region api for SQ")
Fixes: 81a4058e0cd0 ("io_uring: use region api for CQ")
Cc: stable@vger.kernel.org
Signed-off-by: Hengyu Liang <hengyul@cs.unc.edu>
---
This goes on top of io_uring-7.3 (c746673517c6), as discussed in [2].

Tested on that branch as a user without CAP_IPC_LOCK and the default
8 MiB limit:

                                                  before  after
  program above                                       21     64
  the same with io_uring_queue_init()                 64     64
  30 processes x 142 NO_MMAP rings of 64 entries       7     30

The last row is what 30 PostgreSQL 18 clusters with default settings
create. The number is how many of them got all their rings.

The per-user locked_vm count stays balanced over ring create, close and
resize. Provided buffer rings and IORING_REGISTER_MEM_REGION regions on
user memory are charged as before. Every liburing test gives the same
result with and without the patch.

Not changed here: provided buffer rings on user memory, which is what
io_uring_setup_buf_ring() registers, were not charged in v6.13 either
(64 of 64 rings of 32768 entries then, 16 now). I can send a patch for
those as well if you want them handled the same way.

[2] https://lore.kernel.org/io-uring/b5a33433-b0b9-4231-9998-23e2a2202091@kernel.dk/

 io_uring/io_uring.c |  8 ++++----
 io_uring/kbuf.c     |  2 +-
 io_uring/memmap.c   |  7 ++++---
 io_uring/memmap.h   |  2 +-
 io_uring/register.c | 10 +++++-----
 io_uring/zcrx.c     |  2 +-
 6 files changed, 16 insertions(+), 15 deletions(-)

diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c
index c2ce83c7c1f1..2a16369894d4 100644
--- a/io_uring/io_uring.c
+++ b/io_uring/io_uring.c
@@ -2070,8 +2070,8 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned int nr)
 
 static void io_rings_free(struct io_ring_ctx *ctx)
 {
-	io_free_region(ctx->user, &ctx->sq_region);
-	io_free_region(ctx->user, &ctx->ring_region);
+	io_free_region(NULL, &ctx->sq_region);
+	io_free_region(NULL, &ctx->ring_region);
 	ctx->rings = NULL;
 	RCU_INIT_POINTER(ctx->rings_rcu, NULL);
 	ctx->sq_sqes = NULL;
@@ -2735,7 +2735,7 @@ static __cold int io_allocate_scq_urings(struct io_ring_ctx *ctx,
 		rd.user_addr = p->cq_off.user_addr;
 		rd.flags |= IORING_MEM_REGION_TYPE_USER;
 	}
-	ret = io_create_region(ctx, &ctx->ring_region, &rd, IORING_OFF_CQ_RING);
+	ret = io_create_region(NULL, &ctx->ring_region, &rd, IORING_OFF_CQ_RING);
 	if (ret)
 		return ret;
 	ctx->rings = rings = io_region_get_ptr(&ctx->ring_region);
@@ -2749,7 +2749,7 @@ static __cold int io_allocate_scq_urings(struct io_ring_ctx *ctx,
 		rd.user_addr = p->sq_off.user_addr;
 		rd.flags |= IORING_MEM_REGION_TYPE_USER;
 	}
-	ret = io_create_region(ctx, &ctx->sq_region, &rd, IORING_OFF_SQES);
+	ret = io_create_region(NULL, &ctx->sq_region, &rd, IORING_OFF_SQES);
 	if (ret) {
 		io_rings_free(ctx);
 		return ret;
diff --git a/io_uring/kbuf.c b/io_uring/kbuf.c
index 7c309173dd19..7c59ab9cfc8b 100644
--- a/io_uring/kbuf.c
+++ b/io_uring/kbuf.c
@@ -676,7 +676,7 @@ int io_register_pbuf_ring(struct io_ring_ctx *ctx, void __user *arg)
 		rd.user_addr = reg.ring_addr;
 		rd.flags |= IORING_MEM_REGION_TYPE_USER;
 	}
-	ret = io_create_region(ctx, &bl->region, &rd, mmap_offset);
+	ret = io_create_region(ctx->user, &bl->region, &rd, mmap_offset);
 	if (ret)
 		goto fail;
 	br = io_region_get_ptr(&bl->region);
diff --git a/io_uring/memmap.c b/io_uring/memmap.c
index da2328b52b38..51afe5d40b1d 100644
--- a/io_uring/memmap.c
+++ b/io_uring/memmap.c
@@ -192,7 +192,7 @@ static int io_region_allocate_pages(struct io_mapped_region *mr,
 	return 0;
 }
 
-int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
+int io_create_region(struct user_struct *user, struct io_mapped_region *mr,
 		     struct io_uring_region_desc *reg,
 		     unsigned long mmap_offset)
 {
@@ -223,9 +223,10 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
 	/*
 	 * Only pinned user memory counts against RLIMIT_MEMLOCK, kernel
 	 * allocated regions are memcg accounted through GFP_KERNEL_ACCOUNT.
+	 * The SQ/CQ rings are never charged, their callers pass a NULL user.
 	 */
 	if (reg->flags & IORING_MEM_REGION_TYPE_USER)
-		ret = io_region_pin_pages(mr, reg, ctx->user);
+		ret = io_region_pin_pages(mr, reg, user);
 	else
 		ret = io_region_allocate_pages(mr, reg, mmap_offset);
 	if (ret)
@@ -236,7 +237,7 @@ int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
 		goto out_free;
 	return 0;
 out_free:
-	io_free_region(ctx->user, mr);
+	io_free_region(user, mr);
 	return ret;
 }
 
diff --git a/io_uring/memmap.h b/io_uring/memmap.h
index f4cfbb6b9a1f..0714bb5a9616 100644
--- a/io_uring/memmap.h
+++ b/io_uring/memmap.h
@@ -18,7 +18,7 @@ unsigned long io_uring_get_unmapped_area(struct file *file, unsigned long addr,
 int io_uring_mmap(struct file *file, struct vm_area_struct *vma);
 
 void io_free_region(struct user_struct *user, struct io_mapped_region *mr);
-int io_create_region(struct io_ring_ctx *ctx, struct io_mapped_region *mr,
+int io_create_region(struct user_struct *user, struct io_mapped_region *mr,
 		     struct io_uring_region_desc *reg,
 		     unsigned long mmap_offset);
 
diff --git a/io_uring/register.c b/io_uring/register.c
index 02bc103bcc9d..d79971c57ace 100644
--- a/io_uring/register.c
+++ b/io_uring/register.c
@@ -480,8 +480,8 @@ struct io_ring_ctx_rings {
 static void io_register_free_rings(struct io_ring_ctx *ctx,
 				   struct io_ring_ctx_rings *r)
 {
-	io_free_region(ctx->user, &r->sq_region);
-	io_free_region(ctx->user, &r->ring_region);
+	io_free_region(NULL, &r->sq_region);
+	io_free_region(NULL, &r->ring_region);
 }
 
 #define swap_old(ctx, o, n, field)		\
@@ -529,7 +529,7 @@ static int io_register_resize_rings(struct io_ring_ctx *ctx, void __user *arg)
 		rd.user_addr = p->cq_off.user_addr;
 		rd.flags |= IORING_MEM_REGION_TYPE_USER;
 	}
-	ret = io_create_region(ctx, &n.ring_region, &rd, IORING_OFF_CQ_RING);
+	ret = io_create_region(NULL, &n.ring_region, &rd, IORING_OFF_CQ_RING);
 	if (ret)
 		return ret;
 
@@ -559,7 +559,7 @@ static int io_register_resize_rings(struct io_ring_ctx *ctx, void __user *arg)
 		rd.user_addr = p->sq_off.user_addr;
 		rd.flags |= IORING_MEM_REGION_TYPE_USER;
 	}
-	ret = io_create_region(ctx, &n.sq_region, &rd, IORING_OFF_SQES);
+	ret = io_create_region(NULL, &n.sq_region, &rd, IORING_OFF_SQES);
 	if (ret) {
 		io_register_free_rings(ctx, &n);
 		return ret;
@@ -730,7 +730,7 @@ static int io_register_mem_region(struct io_ring_ctx *ctx, void __user *uarg)
 	    !(ctx->flags & IORING_SETUP_R_DISABLED))
 		return -EINVAL;
 
-	ret = io_create_region(ctx, &region, &rd, IORING_MAP_OFF_PARAM_REGION);
+	ret = io_create_region(ctx->user, &region, &rd, IORING_MAP_OFF_PARAM_REGION);
 	if (ret)
 		return ret;
 	if (copy_to_user(rd_uptr, &rd, sizeof(rd))) {
diff --git a/io_uring/zcrx.c b/io_uring/zcrx.c
index beb35077f3ce..939885985333 100644
--- a/io_uring/zcrx.c
+++ b/io_uring/zcrx.c
@@ -432,7 +432,7 @@ static int io_allocate_rbuf_ring(struct io_ring_ctx *ctx,
 	mmap_offset = IORING_MAP_OFF_ZCRX_REGION;
 	mmap_offset += (u64)id << IORING_OFF_ZCRX_SHIFT;
 
-	ret = io_create_region(ctx, &ifq->rq_region, rd, mmap_offset);
+	ret = io_create_region(ctx->user, &ifq->rq_region, rd, mmap_offset);
 	if (ret < 0)
 		return ret;
 

base-commit: c746673517c6ce9f5400cc0ea23e10ef5382eddb
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH] io_uring: do not charge user provided SQ/CQ rings to RLIMIT_MEMLOCK
  2026-10-08 17:06       ` [PATCH] io_uring: do not charge user provided SQ/CQ rings " Hengyu Liang
@ 2026-10-08 19:11         ` Jens Axboe
  0 siblings, 0 replies; 7+ messages in thread
From: Jens Axboe @ 2026-10-08 19:11 UTC (permalink / raw)
  To: Hengyu Liang; +Cc: Pavel Begunkov, David Wei, io-uring, linux-kernel, netdev

On 10/8/26 11:06 AM, Hengyu Liang wrote:
> This goes on top of io_uring-7.3 (c746673517c6), as discussed in [2].
> 
> Tested on that branch as a user without CAP_IPC_LOCK and the default
> 8 MiB limit:
> 
>                                                   before  after
>   program above                                       21     64
>   the same with io_uring_queue_init()                 64     64
>   30 processes x 142 NO_MMAP rings of 64 entries       7     30
> 
> The last row is what 30 PostgreSQL 18 clusters with default settings
> create. The number is how many of them got all their rings.
> 
> The per-user locked_vm count stays balanced over ring create, close and
> resize. Provided buffer rings and IORING_REGISTER_MEM_REGION regions on
> user memory are charged as before. Every liburing test gives the same
> result with and without the patch.
> 
> Not changed here: provided buffer rings on user memory, which is what
> io_uring_setup_buf_ring() registers, were not charged in v6.13 either
> (64 of 64 rings of 32768 entries then, 16 now). I can send a patch for
> those as well if you want them handled the same way.

Honestly, after taking a closer look at this, I think we're better off
with your original patch and one on top for pbuf rings. Please check:

https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git/log/?h=io_uring-7.3

for the top 2 commits. If you can re-test one more time, that'd be
great...

-- 
Jens Axboe

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-10-08 19:11 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 21:38 [PATCHSET 0/2] Fix memlock account for kernel backed memory Jens Axboe
2026-10-07 21:38 ` [PATCH 1/2] Revert "io_uring/memmap: account the pages a compound region really uses" Jens Axboe
2026-10-07 21:38 ` [PATCH 2/2] io_uring/memmap: only charge pinned user memory to RLIMIT_MEMLOCK Jens Axboe
2026-10-08  4:57   ` Hengyu Liang
2026-10-08 14:54     ` Jens Axboe
2026-10-08 17:06       ` [PATCH] io_uring: do not charge user provided SQ/CQ rings " Hengyu Liang
2026-10-08 19:11         ` Jens Axboe

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox