public inbox for io-uring@vger.kernel.org
 help / color / mirror / Atom feed
* [REGRESSION] io_uring/futex: scalar wait/wake slowdown after 079afb081c42
@ 2026-07-30 10:42 Chengfeng Lin
  2026-07-30 14:34 ` Jens Axboe
  0 siblings, 1 reply; 4+ messages in thread
From: Chengfeng Lin @ 2026-07-30 10:42 UTC (permalink / raw)
  To: Jens Axboe
  Cc: Pavel Begunkov, Robert Morris, io-uring, linux-kernel,
	regressions

Hi Jens,

I tested 079afb081c42 against its direct parent on bare metal. In a narrow
scalar io_uring futex wait/wake workload, the child was 9.27% slower. A
separate 338-line standalone reproduced the result at 10.02%. All compared
kernels actually ran with preempt=full.

#regzbot introduced: 079afb081c4288e94d5e4223d3eb6306d853c68b
#regzbot title: io_uring scalar futex wait/wake slowdown

This is a focused synthetic microbenchmark, not an application benchmark. It
uses one raw-UAPI ring and 32 cacheline-separated private futex words on one
pinned P-core. Each timed cycle submits 32 scalar IORING_OP_FUTEX_WAIT
requests, then 32 scalar IORING_OP_FUTEX_WAKE requests, and drains exactly 64
CQEs. Every wait must return 0 and every wake must return 1.

I used a fresh boot for each point:

  6a8118a77eec parent A -> 079afb081c42 child -> 6a8118a77eec parent B

Each point had 3 warm-up rounds and 15 measured rounds. Every measured round
ran 512 cycles, or 16,384 wait/wake pairs. The results in ns/pair were:

  implementation   parent A    child    parent B    child vs midpoint
  formal              180.079  196.647     179.856        +9.268%
  standalone          180.171  197.856     179.502       +10.020%

For the formal source, dropping the first measured round gave +9.269%. Parent
drift was -0.124%, and the maximum CV was 0.169%. The standalone drop-first
result was +9.996%, with -0.371% parent drift. All 90 scalar timing rows
passed the CQE, result, timeout, overflow, outstanding-request, and CPU
checks.

An untimed child trace also hit io_futex_prep(), io_futex_wait(),
io_futex_wake(), and io_futex_complete() with the expected request counts.

A matched WAITV -> WAKE profile changed by only +1.385%, below my preregistered
5% signal gate, so my claim is limited to scalar wait/wake.

I understand that 079afb fixes the exit-time use-after-free by keeping pending
private futex waits visible to cancellation before their mm state disappears.
Scalar WAIT and WAKE both use io_futex_prep(), so in the child both sides of
each measured pair execute the added tracking call. I am not suggesting a
revert.

Is this per-request cost an expected trade-off for the lifetime fix, or could
the same exit/mm-lifetime guarantee be retained with cheaper tracking?

Evidence bundle:

  https://github.com/lcf0399/linux-regression-evidence/tree/65c8cbf86f40cbe759e3f7db1d29d152ba03a8f2/io-uring-futex-inflight-wait-wake

Standalone reproducer:

  https://github.com/lcf0399/linux-regression-evidence/tree/65c8cbf86f40cbe759e3f7db1d29d152ba03a8f2/io-uring-futex-inflight-wait-wake/reproducer

Thanks,
Chengfeng

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-07-31  8:31 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-30 10:42 [REGRESSION] io_uring/futex: scalar wait/wake slowdown after 079afb081c42 Chengfeng Lin
2026-07-30 14:34 ` Jens Axboe
2026-07-30 15:49   ` Jens Axboe
2026-07-31  8:31     ` Chengfeng Lin

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox