public inbox for io-uring@vger.kernel.org
 help / color / mirror / Atom feed
* [PATCH] io_uring/cmd_net: prevent infinite retry loop on unextractable timestamp skb
@ 2026-10-08  5:07 Bui Viet Dung
  2026-10-08 14:43 ` Pavel Begunkov
  0 siblings, 1 reply; 3+ messages in thread
From: Bui Viet Dung @ 2026-10-08  5:07 UTC (permalink / raw)
  To: Jens Axboe, Pavel Begunkov
  Cc: Willem de Bruijn, lollipopkit, io-uring, linux-kernel, stable,
	Bui Viet Dung

In io_uring_cmd_timestamp(), the skb processing loop terminates whenever
io_process_timestamp_skb() returns a non-zero value. If the failure is
due to a full CQ ring (-ENOBUFS), the loop stops and returns -ENOBUFS,
ending the multishot command so userspace can drain the CQ.

However, if io_process_timestamp_skb() fails because
skb_get_tx_timestamp() returns a negative error (such as -ENOENT when a
timestamp cannot be extracted due to changed socket options or absent
hardware timestamp data), ret is not -ENOBUFS. The unhandled skb is
then spliced back onto the head of sk->sk_error_queue and the function
returns -EAGAIN.

Because -EAGAIN leaves the multishot apoll armed on EPOLLERR, and the
unextractable skb remains in sk_error_queue asserting EPOLLERR,
io_uring_cmd_timestamp() is immediately re-invoked on the same skb.
This results in an infinite busy-loop consuming 100% CPU and completely
blocking progress on any subsequent valid timestamp packets queued behind
it.

Only break out of the processing loop when CQ space is exhausted
(ret == -ENOBUFS). For skbs where timestamp extraction fails, dequeue and
consume the invalid skb matching the behavior of sock_recv_errqueue(),
allowing the queue to make forward progress.

The issue was discovered via manual code audit of io_uring/cmd_net.c
and review of the error handling paths in TX_TIMESTAMP command.

Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
Cc: stable@vger.kernel.org
Signed-off-by: Bui Viet Dung <dungvn2345@gmail.com>
---
 io_uring/cmd_net.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
index 90d4ec7cc761..a5a1b8b66b1f 100644
--- a/io_uring/cmd_net.c
+++ b/io_uring/cmd_net.c
@@ -138,7 +138,7 @@ static int io_uring_cmd_timestamp(struct socket *sock,
 		if (!skb)
 			break;
 		ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
-		if (ret)
+		if (ret == -ENOBUFS)
 			break;
 		__skb_dequeue(&list);
 		consume_skb(skb);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH] io_uring/cmd_net: prevent infinite retry loop on unextractable timestamp skb
  2026-10-08  5:07 [PATCH] io_uring/cmd_net: prevent infinite retry loop on unextractable timestamp skb Bui Viet Dung
@ 2026-10-08 14:43 ` Pavel Begunkov
  2026-10-08 14:59   ` Bui Viet Dung
  0 siblings, 1 reply; 3+ messages in thread
From: Pavel Begunkov @ 2026-10-08 14:43 UTC (permalink / raw)
  To: Bui Viet Dung, Jens Axboe
  Cc: Willem de Bruijn, lollipopkit, io-uring, linux-kernel, stable

On 10/8/26 06:07, Bui Viet Dung wrote:
> In io_uring_cmd_timestamp(), the skb processing loop terminates whenever
> io_process_timestamp_skb() returns a non-zero value. If the failure is
> due to a full CQ ring (-ENOBUFS), the loop stops and returns -ENOBUFS,
> ending the multishot command so userspace can drain the CQ.
> 
> However, if io_process_timestamp_skb() fails because
> skb_get_tx_timestamp() returns a negative error (such as -ENOENT when a
> timestamp cannot be extracted due to changed socket options or absent
> hardware timestamp data), ret is not -ENOBUFS. The unhandled skb is
> then spliced back onto the head of sk->sk_error_queue and the function
> returns -EAGAIN.
> 
> Because -EAGAIN leaves the multishot apoll armed on EPOLLERR, and the
> unextractable skb remains in sk_error_queue asserting EPOLLERR,
> io_uring_cmd_timestamp() is immediately re-invoked on the same skb.
> This results in an infinite busy-loop consuming 100% CPU and completely
> blocking progress on any subsequent valid timestamp packets queued behind
> it.
> 
> Only break out of the processing loop when CQ space is exhausted
> (ret == -ENOBUFS). For skbs where timestamp extraction fails, dequeue and
> consume the invalid skb matching the behavior of sock_recv_errqueue(),
> allowing the queue to make forward progress.
> 
> The issue was discovered via manual code audit of io_uring/cmd_net.c
> and review of the error handling paths in TX_TIMESTAMP command.
> 
> Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
> Cc: stable@vger.kernel.org
> Signed-off-by: Bui Viet Dung <dungvn2345@gmail.com>
> ---
>   io_uring/cmd_net.c | 2 +-
>   1 file changed, 1 insertion(+), 1 deletion(-)
> 
> diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
> index 90d4ec7cc761..a5a1b8b66b1f 100644
> --- a/io_uring/cmd_net.c
> +++ b/io_uring/cmd_net.c
> @@ -138,7 +138,7 @@ static int io_uring_cmd_timestamp(struct socket *sock,
>   		if (!skb)
>   			break;
>   		ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
> -		if (ret)
> +		if (ret == -ENOBUFS)
>   			break;

Sounds fine since there are only timestamp skbs in this list, do
you have a test case?

-- 
Pavel Begunkov


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] io_uring/cmd_net: prevent infinite retry loop on unextractable timestamp skb
  2026-10-08 14:43 ` Pavel Begunkov
@ 2026-10-08 14:59   ` Bui Viet Dung
  0 siblings, 0 replies; 3+ messages in thread
From: Bui Viet Dung @ 2026-10-08 14:59 UTC (permalink / raw)
  To: Pavel Begunkov, Jens Axboe
  Cc: Willem de Bruijn, lollipopkit, io-uring, linux-kernel, stable

On Thu, Oct 8, 2026 at 3:43 PM Pavel Begunkov <asml.silence@gmail.com> wrote:
> Sounds fine since there are only timestamp skbs in this list, do
> you have a test case?

Hi Pavel,

Thanks for the review.

The issue was spotted during code inspection of the error handling path
in io_uring_cmd_timestamp(): specifically when skb_get_tx_timestamp() returns
a negative error (e.g. -ENOENT due to missing PHC bindings with
SOF_TIMESTAMPING_BIND_PHC or unavailable timestamp data). Since the
multishot apoll remains armed on EPOLLERR and the unhandled skb is spliced
back onto the head of sk_error_queue, it immediately re-triggers the command
in a tight busy-loop.

Below is a standalone test program demonstrating the TX_TIMESTAMP command
setup and execution with CQE32:

--- test_ts_cmd.c ---
#define _GNU_SOURCE
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include <sys/socket.h>
#include <netinet/in.h>
#include <arpa/inet.h>
#include <linux/net_tstamp.h>
#include <linux/io_uring.h>
#include <sys/syscall.h>
#include <sys/mman.h>

#define IORING_OP_URING_CMD 46
#define SOCKET_URING_OP_TX_TIMESTAMP 4
#define IORING_SETUP_CQE32 (1U << 11)
#define SOF_TIMESTAMPING_OPT_TSONLY (1 << 11)

int main(void)
{
	int s, r_sock, ring_fd, val;
	struct sockaddr_in r_sin;
	socklen_t r_len = sizeof(r_sin);
	struct io_uring_params p;
	void *sq_ptr, *cq_ptr;
	struct io_uring_sqe *sqes;
	struct io_uring_cqe *cqes;
	unsigned *sq_tail, *sq_array, *cq_head, *cq_tail;
	char dummy = 'A';

	/* Receiver socket */
	r_sock = socket(AF_INET, SOCK_DGRAM, 0);
	memset(&r_sin, 0, sizeof(r_sin));
	r_sin.sin_family = AF_INET;
	r_sin.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
	bind(r_sock, (struct sockaddr *)&r_sin, sizeof(r_sin));
	getsockname(r_sock, (struct sockaddr *)&r_sin, &r_len);

	/* Sender socket with TX software timestamp */
	s = socket(AF_INET, SOCK_DGRAM, 0);
	val = SOF_TIMESTAMPING_SOFTWARE | SOF_TIMESTAMPING_TX_SOFTWARE |
	      SOF_TIMESTAMPING_OPT_ID | SOF_TIMESTAMPING_OPT_TSONLY;
	setsockopt(s, SOL_SOCKET, SO_TIMESTAMPING, &val, sizeof(val));

	/* Send packet to queue timestamp skb into sk_error_queue */
	sendto(s, &dummy, sizeof(dummy), 0, (struct sockaddr *)&r_sin, sizeof(r_sin));
	usleep(10000);

	/* Initialize io_uring with CQE32 */
	memset(&p, 0, sizeof(p));
	p.flags = IORING_SETUP_CQE32;
	ring_fd = syscall(__NR_io_uring_setup, 4, &p);
	if (ring_fd < 0)
		return 1;

	sq_ptr = mmap(NULL, p.sq_off.array + p.sq_entries * sizeof(unsigned),
		      PROT_READ | PROT_WRITE, MAP_SHARED | MAP_POPULATE,
		      ring_fd, IORING_OFF_SQ_RING);
	sqes = mmap(NULL, p.sq_entries * sizeof(struct io_uring_sqe),
		    PROT_READ | PROT_WRITE, MAP_SHARED | MAP_POPULATE,
		    ring_fd, IORING_OFF_SQES);
	cq_ptr = mmap(NULL, p.cq_off.cqes + p.cq_entries * 2 * sizeof(struct io_uring_cqe),
		      PROT_READ | PROT_WRITE, MAP_SHARED | MAP_POPULATE,
		      ring_fd, IORING_OFF_CQ_RING);

	sq_tail = (unsigned *)(sq_ptr + p.sq_off.tail);
	sq_array = (unsigned *)(sq_ptr + p.sq_off.array);
	cq_head = (unsigned *)(cq_ptr + p.cq_off.head);
	cq_tail = (unsigned *)(cq_ptr + p.cq_off.tail);

	memset(&sqes[0], 0, sizeof(struct io_uring_sqe));
	sqes[0].opcode = IORING_OP_URING_CMD;
	sqes[0].fd = s;
	sqes[0].cmd_op = SOCKET_URING_OP_TX_TIMESTAMP;

	sq_array[0] = 0;
	*sq_tail = 1;

	/* Submit command and wait for CQE */
	syscall(__NR_io_uring_enter, ring_fd, 1, 1, 1, NULL, 0);

	close(ring_fd);
	close(s);
	close(r_sock);
	return 0;
}
---

If desired, I can format this into a test case for liburing under
test/timestamp.c.

Thanks,
Bui Viet Dung

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-08 14:59 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08  5:07 [PATCH] io_uring/cmd_net: prevent infinite retry loop on unextractable timestamp skb Bui Viet Dung
2026-10-08 14:43 ` Pavel Begunkov
2026-10-08 14:59   ` Bui Viet Dung

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox