From: "Björn Töpel" <bjorn@kernel.org>
To: Magnus Karlsson <magnus.karlsson@intel.com>,
Maciej Fijalkowski <maciej.fijalkowski@intel.com>,
Stanislav Fomichev <sdf@fomichev.me>,
"David S. Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@kernel.org>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
Simon Horman <horms@kernel.org>, Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Randy Dunlap <rdunlap@infradead.org>,
Alexander Duyck <alexanderduyck@fb.com>,
kernel-team@meta.com, Andrew Lunn <andrew+netdev@lunn.ch>,
Jesper Dangaard Brouer <hawk@kernel.org>,
Ilias Apalodimas <ilias.apalodimas@linaro.org>,
Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
John Fastabend <john.fastabend@gmail.com>,
Pavel Begunkov <asml.silence@gmail.com>,
Jens Axboe <axboe@kernel.dk>, Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
netdev@vger.kernel.org, bpf@vger.kernel.org,
io-uring@vger.kernel.org
Cc: "Björn Töpel" <bjorn@kernel.org>,
"Mike Marciniszyn (Meta)" <mike.marciniszyn@gmail.com>,
"Weiming Shi" <bestswngs@gmail.com>,
"Nikolay Aleksandrov" <razor@blackwall.org>,
"David Wei" <dw@davidwei.uk>,
"Alexander Lobakin" <aleksander.lobakin@intel.com>,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
"Mina Almasry" <almasrymina@google.com>
Subject: [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers
Date: Fri, 2 Oct 2026 21:00:01 +0200 [thread overview]
Message-ID: <20261002190018.696925-1-bjorn@kernel.org> (raw)
Hi!
Here's an RFC for you all to enjoy with your favorite Friday
$BEVERAGE.
AF_XDP zero copy needs a second receive allocator in every driver.
The driver takes xdp_buff_xsk objects from the XSK buffer pool beside
its page_pool path and owns their lifetime. Drivers built around
page_pool and the queue management API, which already support devmem
and io_uring zero-copy receive, have to duplicate their receive path.
This series makes UMEM a page-pool memory provider instead. Each
aligned 4 KiB chunk is a NET_IOV_XSK net_iov. Binding a zero-copy
socket installs the provider on the queue through the queue
management API. The driver keeps its page_pool allocation, DMA sync,
recycling and refill code. The provider consumes FILL, drops invalid
and repeated addresses, and owns RX need-wakeup.
An XSKMAP redirect publishes the UMEM address when every buffer of the
frame belongs to the socket's provider. XDP_PASS and other redirect
targets copy into page-backed memory first. TX is unchanged; drivers
read socket TX descriptors directly.
Core changes:
- MP_CAP_READABLE and MP_CAP_FRAG. A readable provider may back
header and regular pools, and XDP may run on its buffers. devmem
and io_uring set MP_CAP_FRAG and keep their behavior.
- Providers request RX headroom through the queue configuration.
- Readable net_iov areas. netmem_address() resolves provider memory
with a load and a shift, without an indirect call.
- A refill-completion callback. The XSK provider keeps NAPI scheduled
while FILL has entries, and sets NEED_WAKEUP when FILL is empty.
- Batched page-pool release for objects handed to userspace.
- xdp_buff carries the netmem and a pointer to kernel-owned shared
info, so fragment metadata never lives in user-writable UMEM. It
grows from 56 to 64 bytes on 64-bit.
Differences from classic zero copy:
- 4 KiB pages and aligned 4 KiB chunks only. Other layouts fail to
bind with -EOPNOTSUPP, but can be supported in the future.
- Generic XDP and CPUMAP cannot deliver to a provider-backed socket.
Classic zero copy is unchanged. This does not propose converting
existing drivers; it is for page_pool drivers without zero copy.
fbnic is the only driver user, +640/-122 for RX and TX. I have bnxt
working with AF_XDP, plus some performance patches/fixes for the
AF_XDP core on top of this -- but let's start with these patches.
Patches 1-2 fix bugs in net-next that the series hits. They are
carried here so the series can be tested on its own.
1. "xdp: Size zero-copy skb heads by their contents". XDP_PASS of a
zero-copy buffer copies it into an skb whose head is sized by the
XSK frame size, which leaves no room for skb_shared_info. A frame
that fills its buffer overwrites skb_shared_info. Provider
buffers take this path on XDP_PASS, which the usual AF_XDP
program returns when no socket is bound to the queue. Reproduced
on fbnic in QEMU.
2. "eth: fbnic: Report the logical XDP RX queue". fbnic reports
queue 0 in rx_queue_index for every queue. AF_XDP drops frames
whose queue differs from the socket's, and the usual program
looks up its socket by that index, so zero copy works on queue 0
only.
Feedback wanted on:
0. General thoughts on extending the page pool provider.
1. Does refill_done belong in page_pool, or should finite providers
keep NAPI scheduled some other way?
2. PP_FLAG_ALLOW_UNREADABLE_NETMEM is how a pool picks up any
provider, readable or not. A driver that only wants AF_XDP must
set it, and then passes the core checks for devmem and io_uring
too if it supports header split. Drivers that set the flag today
may also split buffers, for example mlx5 through
page_pool_fragment_netmem(). Only QCFG_RX_HEADROOM keeps the
unsplittable XSK provider away from them. Split the flag, or let
drivers declare support for unsplittable providers?
3. Is growing xdp_buff by 8 bytes acceptable?
4. How should userspace learn the chunk constraints before bind?
This is a way to move code from the drivers to the core, reducing the
work for driver developers.
I hope to see you at LPC next week! Particular the netdev and bpf MC,
and the AF_XDP BoF on Monday.
Björn
Björn Töpel (15):
xdp: Size zero-copy skb heads by their contents
eth: fbnic: Report the logical XDP RX queue
net: Add memory provider capabilities
net: Let memory providers set RX buffer headroom
page_pool: Extend memory provider operations
xdp: Track non-page netmem in receive buffers
xsk: Keep the DMA mapping in the buffer pool
xsk: Handle a detached FILL ring in RX wakeup
xsk: Add a page-pool memory provider for UMEM
xsk: Add RX helpers for page-pool drivers
xdp: Copy provider buffers on pass and redirect
xsk: Receive provider UMEM without copying
eth: fbnic: Support AF_XDP zero-copy receive
eth: fbnic: Support AF_XDP zero-copy transmit
Documentation: xsk: Document page-pool zero copy
Documentation/networking/af_xdp.rst | 65 ++
Documentation/networking/netmem.rst | 7 +-
.../net/ethernet/meta/fbnic/fbnic_ethtool.c | 5 +
.../net/ethernet/meta/fbnic/fbnic_netdev.c | 164 ++++-
.../net/ethernet/meta/fbnic/fbnic_netdev.h | 4 +
drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 561 ++++++++++++---
drivers/net/ethernet/meta/fbnic/fbnic_txrx.h | 24 +
include/net/netdev_queues.h | 6 +
include/net/netmem.h | 32 +-
include/net/page_pool/helpers.h | 62 +-
include/net/page_pool/memory_provider.h | 39 +-
include/net/page_pool/types.h | 8 +-
include/net/xdp.h | 64 +-
include/net/xdp_sock.h | 10 +
include/net/xdp_sock_drv.h | 58 +-
include/net/xsk_buff_pool.h | 7 +-
io_uring/zcrx.c | 4 +-
net/core/dev.c | 4 +-
net/core/dev.h | 5 +
net/core/devmem.c | 7 +-
net/core/filter.c | 28 +-
net/core/netdev_config.c | 2 +
net/core/netdev_rx_queue.c | 61 +-
net/core/page_pool.c | 72 +-
net/core/xdp.c | 112 ++-
net/ethtool/rings.c | 18 +-
net/xdp/Kconfig | 1 +
net/xdp/xsk.c | 253 ++++++-
net/xdp/xsk.h | 58 ++
net/xdp/xsk_buff_pool.c | 667 +++++++++++++++++-
30 files changed, 2164 insertions(+), 244 deletions(-)
base-commit: 071876fd50482a68603a9460d80dd6dd58827ee1
--
2.55.0
next reply other threads:[~2026-10-02 19:00 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 19:00 Björn Töpel [this message]
2026-10-02 19:00 ` [RFC net-next 01/15] xdp: Size zero-copy skb heads by their contents Björn Töpel
2026-10-02 19:00 ` [RFC net-next 02/15] eth: fbnic: Report the logical XDP RX queue Björn Töpel
2026-10-02 19:00 ` [RFC net-next 03/15] net: Add memory provider capabilities Björn Töpel
2026-10-02 19:00 ` [RFC net-next 04/15] net: Let memory providers set RX buffer headroom Björn Töpel
2026-10-02 19:00 ` [RFC net-next 05/15] page_pool: Extend memory provider operations Björn Töpel
2026-10-02 19:00 ` [RFC net-next 06/15] xdp: Track non-page netmem in receive buffers Björn Töpel
2026-10-02 19:00 ` [RFC net-next 07/15] xsk: Keep the DMA mapping in the buffer pool Björn Töpel
2026-10-02 19:00 ` [RFC net-next 08/15] xsk: Handle a detached FILL ring in RX wakeup Björn Töpel
2026-10-02 19:00 ` [RFC net-next 09/15] xsk: Add a page-pool memory provider for UMEM Björn Töpel
2026-10-02 19:00 ` [RFC net-next 10/15] xsk: Add RX helpers for page-pool drivers Björn Töpel
2026-10-02 19:00 ` [RFC net-next 11/15] xdp: Copy provider buffers on pass and redirect Björn Töpel
2026-10-02 19:00 ` [RFC net-next 12/15] xsk: Receive provider UMEM without copying Björn Töpel
2026-10-02 19:00 ` [RFC net-next 13/15] eth: fbnic: Support AF_XDP zero-copy receive Björn Töpel
2026-10-02 19:00 ` [RFC net-next 14/15] eth: fbnic: Support AF_XDP zero-copy transmit Björn Töpel
2026-10-02 19:00 ` [RFC net-next 15/15] Documentation: xsk: Document page-pool zero copy Björn Töpel
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002190018.696925-1-bjorn@kernel.org \
--to=bjorn@kernel.org \
--cc=aleksander.lobakin@intel.com \
--cc=alexanderduyck@fb.com \
--cc=almasrymina@google.com \
--cc=andrew+netdev@lunn.ch \
--cc=andrii@kernel.org \
--cc=asml.silence@gmail.com \
--cc=ast@kernel.org \
--cc=axboe@kernel.dk \
--cc=bestswngs@gmail.com \
--cc=bpf@vger.kernel.org \
--cc=corbet@lwn.net \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=dw@davidwei.uk \
--cc=eddyz87@gmail.com \
--cc=edumazet@kernel.org \
--cc=emil@etsalapatis.com \
--cc=hawk@kernel.org \
--cc=horms@kernel.org \
--cc=ihor.solodrai@linux.dev \
--cc=ilias.apalodimas@linaro.org \
--cc=io-uring@vger.kernel.org \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=kernel-team@meta.com \
--cc=kuba@kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=maciej.fijalkowski@intel.com \
--cc=magnus.karlsson@intel.com \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=mike.marciniszyn@gmail.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=razor@blackwall.org \
--cc=rdunlap@infradead.org \
--cc=sdf@fomichev.me \
--cc=skhan@linuxfoundation.org \
--cc=song@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox