public inbox for io-uring@vger.kernel.org
 help / color / mirror / Atom feed
* [SECURITY] io_uring: provided-buffer ownership loss across MSG_WAITALL retry
@ 2026-09-16 21:35 Andres Berbescu
  2026-09-17  4:22 ` Jens Axboe
  0 siblings, 1 reply; 2+ messages in thread
From: Andres Berbescu @ 2026-09-16 21:35 UTC (permalink / raw)
  To: axboe; +Cc: io-uring, sashal, gregkh, security


[-- Attachment #1.1: Type: text/plain, Size: 3254 bytes --]

Hello,

I am reporting one io_uring provided-buffer ownership/accounting bug with
two MSG_WAITALL manifestations: bundle-tail reuse and incremental-buffer
reuse. The two cases violate the same invariant and should be handled as
one report.

Invariant and root cause

No provided-buffer region may be published or allocated to request B while
request A retains an iterator that can later write to it.

During a positive partial MSG_WAITALL receive, io_net_kbuf_recyle() commits
only the current prefix and consumes REQ_F_BUFFERS_COMMIT while the
persistent iterator still names the remainder. Retry continues without
reselection, and terminal completion has no remaining buffer-list ownership
token. In the incremental case, the retained-buffer result from commit is
also discarded.

Consequently, request B can be allocated a region that request A can still
write during its later retry.

Affected range

The bundle feature begins at v6.10 (2f9c9515bdfd), and incremental pbuf
consumption begins at v6.12 (ae98dbf43d75). The submission-relevant
residual begins upstream at v6.17 (41b70df5b38) and in stable 6.12 at
v6.12.44 (fe9da1812f86). It remains present through stable v6.12.110 and
tested mainline f6e7b42bf05b2427fb8a7a1d1c387a86638bb413.

Validation

Both manifestations reproduced 20/20 on each of these unmodified targets:

   - Debian 6.12.107
   - upstream stable 6.12.110
   - mainline f6e7b42bf05b2427fb8a7a1d1c387a86638bb413

Kernel-side tracing records request identity, group, bid, head/tail,
iterator range, commit token and CQE publication. It proves that A still
has write capability when the same physical range is assigned to B. A
minimal diagnostic patch removed both manifestations in 40/40 runs.
Different-group and single-request controls removed the collision in 20/20
runs each.

All runs used UID/EUID 1000, group 1000 only, CapEff=0, and no namespace,
capability, SQPOLL, module or privileged socket option.

Observed impact

The demonstrated effect is deterministic integrity damage inside the
caller's registered userspace pbuf pool. In bundle mode, A can overwrite
4096 bytes assigned to B or produce a 100-byte-A/3996-byte-B mixed record.
In incremental mode, B overwrites 40 bytes already delivered by A.

I am not claiming kernel-memory corruption, arbitrary write, cross-mm or
cross-principal access, confidentiality loss, availability loss, or
privilege escalation.

Duplicate assessment

This is an incomplete-fix/variant of CVE-2025-38730, not a new defect
class. It is code-level distinct from CVE-2026-53191; the CVE-2026-53191
fix is present in every kernel where this report reproduces and does not
repair the ownership transition described above.

Reproducer and evidence

A combined standalone package contains both source reproducers and causal
controls. Complete raw evidence, manifests, kernel traces, the comparison
with CVE-2026-53191, and the diagnostic patch are available privately on
request.

AI assistance was used during analysis and report preparation. In
accordance with the current kernel reporting guidance, I am treating the
issue as public and am not publishing the reproducers.

I will attach full report, please if you consider that full PoC + evidence
is needed just let me know. Thanks

[-- Attachment #1.2: Type: text/html, Size: 3482 bytes --]

[-- Attachment #2: report.md --]
[-- Type: text/markdown, Size: 9543 bytes --]

# KRN-2026-001/002 — io_uring `MSG_WAITALL` retry loses provided-buffer ownership

Track: `io_uring`

## Summary

`io_uring` can publish or reallocate a provided-buffer range while an earlier
`IORING_OP_RECV` still has a persistent iterator capable of writing it. The
trigger is a positive partial `MSG_WAITALL` receive. The retry helper commits
only that internal receive's prefix, consumes the request's one-shot
`REQ_F_BUFFERS_COMMIT` token, and returns `-EAGAIN`. The retry reuses the
existing iterator without reselecting, and terminal completion has no buffer
list/token with which to reconcile the remaining bytes.

This produces two manifestations of one root cause:

1. With a non-incremental bundle, the short prefix advances the head over only
   the first physical entry. The unfilled mapped tail becomes selectable by
   request B while request A can still write it.
2. With an incremental entry, only the first internal receive advances the
   descriptor. A completes later without charging subsequent bytes or setting
   `IORING_CQE_F_BUF_MORE`; request B is then allocated a range whose bytes A's
   CQE already delivered.

Kernel tracing directly proves A's live iterator and B's allocation name the
same addresses. On unmodified Debian 6.12.107, stable 6.12.110 and current
mainline `f6e7b42bf05b`, each manifestation reproduced 20/20. A minimal
diagnostic patch that reconciles the complete selected extent at the retry
transition removed both in 40/40 runs.

Classification: **INCOMPLETE-FIX / VARIANT of CVE-2025-38730**. It is not a
duplicate of CVE-2026-53191, and KRN-2026-001/002 should receive one coordinated
submission rather than two CVEs.

## Security impact

The demonstrated primitive is deterministic integrity corruption inside the
ring owner's registered userspace buffer pool:

- bundle: B completes into a 4096-byte entry and publishes its bid; after that
  CQE, A overwrites all 4096 B bytes. A reverse-completion case leaves one
  CQE-named entry containing 100 bytes from A and 3996 from B;
- incremental: A completes 100 bytes after internal receives of 60 and 40;
  B is allocated at offset 60 and overwrites exactly A's delivered bytes
  60..99.

No kernel-memory write, out-of-bounds kernel access, arbitrary write,
cross-`mm` or cross-principal mutation, confidentiality loss, availability
loss, privilege escalation or LPE is demonstrated or claimed. The selected
iterators remain userspace iterators over memory registered/chosen by the same
ring owner.

## Reachability and preconditions

- ordinary local UID/EUID 1000, `CapEff=0`, no privileged groups;
- `CONFIG_IO_URING=y`, io_uring enabled (`kernel.io_uring_disabled=0`);
- registered provided-buffer ring and `IOSQE_BUFFER_SELECT`;
- `IORING_OP_RECV` on `AF_UNIX` `SOCK_STREAM` with `MSG_WAITALL`;
- for KRN-001: `IORING_RECVSEND_BUNDLE`, non-incremental entries and at least
  two mapped entries;
- for KRN-002: one incremental (`IOU_PBUF_RING_INC`) entry retained by its
  `min_left` policy;
- a positive partial receive followed by a retry, plus a later receive on the
  same group.

No namespace, capability, kernel module, device, privileged socket option,
registered fixed buffer, SQPOLL or race is required.

## Root cause

On stable 6.12.110:

1. `io_recv()` imports provided buffers and a persistent iterator
   (`io_uring/net.c:1195-1204`).
2. A positive result smaller than the `MSG_WAITALL` minimum enters the retry
   path (`net.c:1209-1224`).
3. `io_net_kbuf_recyle()` sets `REQ_F_BL_NO_RECYCLE` and commits only `len`, the
   current internal receive's bytes (`net.c:503-510`).
4. `io_kbuf_commit()` clears `REQ_F_BUFFERS_COMMIT` before advancing the head
   or incremental descriptor (`io_uring/kbuf.c:62-75`). This token cannot be
   used again.
5. `REQ_F_BUFFER_RING` prevents re-selection and `REQ_F_BL_NO_RECYCLE` prevents
   rollback (`io_uring/kbuf.h:134-149`). The iterator nevertheless persists.
6. The next issue begins with `sel.buf_list=NULL`; terminal
   `__io_put_kbuf_ring()` therefore skips commit (`kbuf.c:380-410`).
7. For incremental entries, the false return from `io_kbuf_inc_commit()` means
   the entry was retained, but the retry helper discards it and does not latch
   `REQ_F_BUF_MORE` (`kbuf.c:35-59`).

The violated invariant is:

> No provided-buffer range may be published or allocated to request B while
> request A retains an iterator that can subsequently write that range.

## Kernel-side ownership proof

### Bundle manifestation

```text
A select: head=0, iov[0]=base/4096, iov[1]=base+4096/4096
A short result: len=100, iterator_remaining=8092, commit=1
A commit: head 0->1, commit=0
B allocation/commit: addr=base+4096, head 1->2
B CQE: res=4096, bid=1
A terminal CQE: res=8192, bl=NULL, commit=0
```

Canaries show entry 1 changing from zero, to 4096 `B` bytes after B's CQE, to
4096 `A` bytes after A resumes.

### Incremental manifestation

```text
A select: bid=0, head=0, descriptor=base/8192, requested=100
A short result: len=60, iterator_remaining=40, commit=1
A commit: head 0->0, descriptor=base+60/8132, return=false, commit=0
A CQE: res=100, cflags=0x1 (BUF_MORE absent), bl=NULL
B select: bid=0, head=0, addr=base+60, requested=100
B commit: descriptor=base+160/8032
B CQE: res=100, cflags=0x11
```

Canaries show `[0,100)=A` at A's CQE and `[60,160)=B` at B's CQE, proving the
40-byte post-CQE overwrite.

## Reproduction matrix

| Kernel                                | Bundle | Incremental | Verdict                          |
| ------------------------------------- | ------:| -----------:| -------------------------------- |
| Debian `6.12.107+deb13-cloud-amd64`   | 20/20  | 20/20       | VULNERABLE                       |
| stable `6.12.110`                     | 20/20  | 20/20       | VULNERABLE                       |
| mainline `f6e7b42bf05b` (`7.3.0-rc3`) | 20/20  | 20/20       | VULNERABLE                       |
| stable `6.12.110` + diagnostic patch  | 0/20   | 0/20        | NOT VULNERABLE to tested trigger |

Each run used a fresh Debian KVM boot and recorded `id`, `uname`, OS release,
io_uring sysctl, relevant kernel config, `CapEff`, boot ID, full serial output,
dmesg, commands, runtime, PoC output and hashes. All normal runs were UID 1000
with zero effective capabilities.

Additional negative controls remove, one at a time, `MSG_WAITALL`, bundle mode,
incremental mode, provided-buffer selection, partial completion, shared group,
and request B. The separate-group and single-request controls pass 20/20 each.
The final clean-guest `startup.sh` and `exploit.sh` package runs exit zero.

## Duplicate analysis

### CVE-2025-38730

Upstream commit `41b70df5b38bc80967d2e0ed55cc3c3896bba781` added the
partial-buffer commit on retry. It fixes the original zero-commit spelling but
commits only the current prefix and consumes the only token. The stable 6.12
backport `fe9da1812f8697a38f7e30991d568ec199e16059` is present in every tested
6.12 kernel. This report is therefore an incomplete-fix residual with direct
code and dynamic evidence.

### CVE-2026-53191

Upstream `ed46f39c47eb5530a9c161481a2080d3a869cfaf` and stable 6.12
`f40570fda3f3a1f96aeaa4aef665ba274b2810b5` preserve
`IORING_CQE_F_BUF_MORE` across internal bundle retries. They do not modify the
ownership transition above. Stable 6.12.107/110 and mainline contain that fix
but reproduce both PoCs. Detailed comparison is in the evidence package.

## Version range

- Bundle receive entered in `2f9c9515bdfde9e4df1f35782284074d3625ff8a`
  (first release v6.10). The broad ownership bug family affects v6.10 through
  current; before the 2025 fix it is the original CVE behavior.
- Incremental pbuf consumption entered in
  `ae98dbf43d755b4e111fcd086e53939bef3e9a1a` (first release v6.12) and is
  affected through current.
- The specific submission-relevant residual begins upstream with the partial
  retry fix (first release v6.17) and in stable 6.12 at v6.12.44.
- It is confirmed through stable v6.12.110 and mainline
  `v7.3-rc3-64-gf6e7b42bf05b`.
- No fixed upstream or stable release was found. Pre-feature releases are
  non-affected for the corresponding manifestation; pre-v6.17 feature-bearing
  releases are not safe because they contain the original CVE condition.

## Diagnostic fix experiment

The preserved diagnostic patch changes the retry charge from `len` to
`len + iov_iter_count(iterator)` and propagates the incremental commit result
into `REQ_F_BUF_MORE`. It changes exactly the suspected transition and removes
both PoCs. It is not proposed as the final upstream patch: maintainers should
validate integer width, partial mappings, cancellation and API semantics before
choosing between full reservation, retry re-selection, or whole-extent commit.

## Reproduction

Inside a disposable Debian KVM guest as the ordinary user:

```sh
unzip poc.zip -d krn001002
cd krn001002
./startup.sh
./exploit.sh
```

`startup.sh` builds both preserved harnesses. `exploit.sh` runs both positive
oracles and their packaged controls and exits nonzero on any missing PASS,
stderr output, setup failure or oracle mismatch. It refuses execution outside a
VM and records preconditions and raw logs.

## CVSS v4.0

**4.8 Medium**

```text
CVSS:4.0/AV:L/AC:L/AT:N/PR:L/UI:N/VC:N/VI:L/VA:N/SC:N/SI:N/SA:N
```

Rationale: local access and low privileges are required; the deterministic
trigger has no race prerequisite; the only demonstrated effect is low integrity
loss in the attacker's/ring owner's supplied userspace pool. No subsequent
system impact is demonstrated.

# 

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [SECURITY] io_uring: provided-buffer ownership loss across MSG_WAITALL retry
  2026-09-16 21:35 [SECURITY] io_uring: provided-buffer ownership loss across MSG_WAITALL retry Andres Berbescu
@ 2026-09-17  4:22 ` Jens Axboe
  0 siblings, 0 replies; 2+ messages in thread
From: Jens Axboe @ 2026-09-17  4:22 UTC (permalink / raw)
  To: Andres Berbescu; +Cc: io-uring, sashal, gregkh, security

On 9/16/26 3:35 PM, Andres Berbescu wrote:
> Hello,
> 
> I am reporting one io_uring provided-buffer ownership/accounting bug
> with two MSG_WAITALL manifestations: bundle-tail reuse and
> incremental-buffer reuse. The two cases violate the same invariant and
> should be handled as one report.

For the three you sent:

1) No point CC'ing security@ when you're sending it to a public list as
   well.
2) Send a patch
3) Potential data corruption is a not a security issue in the first
   place.

Everybody has an LLM these days, if you have it find issues or potential
issues in kernel code, have it generate a patch as well and send it
along. The documentation is pretty clear on that. That's more useful
that 20kb+ of drivel report.

-- 
Jens Axboe

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-17  4:22 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 21:35 [SECURITY] io_uring: provided-buffer ownership loss across MSG_WAITALL retry Andres Berbescu
2026-09-17  4:22 ` Jens Axboe

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox