* [RFC net-next 01/15] xdp: Size zero-copy skb heads by their contents
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 02/15] eth: fbnic: Report the logical XDP RX queue Björn Töpel
` (13 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
xdp_build_skb_from_zc() sizes the skb head by the XSK frame size,
which leaves no room for skb_shared_info. With 2 KiB chunks, a
1514 byte frame overwrites 48 bytes of it. Both copies also round
the length up to LARGEST_ALIGN, past the end of the new buffers.
Size the head from headroom plus data, copy exactly the data, and
drop frames whose head does not fit in a page.
Discovered by an AI code review agent. Reproduced on fbnic in QEMU
with AF_XDP zero-copy from the page-pool series: an XDP program that
adds metadata, grows the frame to its end and passes it gets UMEM
bytes in the skb's shared info. A new xskxceiver test,
XDP_PASS_FULL_FRAME, checks this and passes with the fix. The test
will be posted separately.
Fixes: 560d958c6c68 ("xsk: add generic XSk &xdp_buff -> skb conversion")
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
net/core/xdp.c | 14 ++++++++++----
1 file changed, 10 insertions(+), 4 deletions(-)
diff --git a/net/core/xdp.c b/net/core/xdp.c
index 1d679e8fd649..386240bd24c9 100644
--- a/net/core/xdp.c
+++ b/net/core/xdp.c
@@ -709,7 +709,7 @@ static noinline bool xdp_copy_frags_from_zc(struct sk_buff *skb,
}
memcpy(page_address(page) + offset, skb_frag_address(frag),
- LARGEST_ALIGN(len));
+ len);
__skb_fill_page_desc_noacc(sinfo, i, page, offset, len);
tsize += truesize;
@@ -738,17 +738,23 @@ static noinline bool xdp_copy_frags_from_zc(struct sk_buff *skb,
*/
struct sk_buff *xdp_build_skb_from_zc(struct xdp_buff *xdp)
{
+ u32 headroom = xdp->data_meta - xdp->data_hard_start;
const struct xdp_rxq_info *rxq = xdp->rxq;
u32 len = xdp->data_end - xdp->data_meta;
- u32 truesize = xdp->frame_sz;
struct sk_buff *skb = NULL;
struct page_pool *pp;
+ u32 truesize;
int metalen;
void *data;
if (!IS_ENABLED(CONFIG_PAGE_POOL))
return NULL;
+ /* The XSK frame size leaves no room for skb_shared_info. */
+ truesize = SKB_HEAD_ALIGN(headroom + len);
+ if (unlikely(truesize > PAGE_SIZE))
+ return NULL;
+
local_lock_nested_bh(&system_page_pool.bh_lock);
pp = this_cpu_read(system_page_pool.pool);
data = page_pool_dev_alloc_va(pp, &truesize);
@@ -762,9 +768,9 @@ struct sk_buff *xdp_build_skb_from_zc(struct xdp_buff *xdp)
}
skb_mark_for_recycle(skb);
- skb_reserve(skb, xdp->data_meta - xdp->data_hard_start);
+ skb_reserve(skb, headroom);
- memcpy(__skb_put(skb, len), xdp->data_meta, LARGEST_ALIGN(len));
+ memcpy(__skb_put(skb, len), xdp->data_meta, len);
metalen = xdp->data - xdp->data_meta;
if (metalen > 0) {
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 02/15] eth: fbnic: Report the logical XDP RX queue
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
2026-10-02 19:00 ` [RFC net-next 01/15] xdp: Size zero-copy skb heads by their contents Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 03/15] net: Add memory provider capabilities Björn Töpel
` (12 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
XDP gives the RX queue to programs in ctx->rx_queue_index. fbnic
fills it from the header buffer queue index. The header and payload
buffer queues do not carry the netdev queue id, and the header queue
index is zero. So XDP programs see queue 0 for packets on every
queue.
Programs that act on the RX queue get the wrong value. For example,
the default AF_XDP program looks up its socket with
ctx->rx_queue_index, so it cannot find sockets on other queues.
Use the completion queue index. It is the netdev RX queue id, which
queue steering and AF_XDP socket binding use.
Tested on fbnic hardware (50G, firmware 25.08.05-004) with AF_XDP
zero-copy from the page-pool series: sockets on queue 43, and on four
queues at once, receive their packets.
Fixes: 894d4a4ea6cb ("eth: fbnic: move xdp_rxq_info_reg() to resource alloc")
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
index c62d885a678a..0c7b1ea002d8 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
@@ -2072,7 +2072,7 @@ static int fbnic_alloc_rx_qt_resources(struct fbnic_net *fbn,
fbnic_bdq_set_page_size(&qt->sub1, rx_page_size);
- err = xdp_rxq_info_reg(&qt->xdp_rxq, fbn->netdev, qt->sub0.q_idx,
+ err = xdp_rxq_info_reg(&qt->xdp_rxq, fbn->netdev, qt->cmpl.q_idx,
nv->napi.napi_id);
if (err)
goto free_page_pools;
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 03/15] net: Add memory provider capabilities
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
2026-10-02 19:00 ` [RFC net-next 01/15] xdp: Size zero-copy skb heads by their contents Björn Töpel
2026-10-02 19:00 ` [RFC net-next 02/15] eth: fbnic: Report the logical XDP RX queue Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-03 4:13 ` Mina Almasry
2026-10-02 19:00 ` [RFC net-next 04/15] net: Let memory providers set RX buffer headroom Björn Töpel
` (11 subsequent siblings)
14 siblings, 1 reply; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
The current memory providers, devmem and io_uring zero-copy RX, give
memory that the CPU cannot read. The core allows them only on queues
with header-data split on, a zero header-split threshold, and no XDP
program. AF_XDP needs a provider that the CPU can read and that
works with XDP.
Add a capability field to the provider operations, with
MP_CAP_READABLE as the first flag. Apply the header-split and XDP
rules only to providers without it. A queue with an AF_XDP buffer
pool still rejects other providers, but it accepts the provider made
from that same pool.
Provider memory can become free while a queue is restarted, and no
device event reports it. After a restart, and after a failed restart
is rolled back, schedule the queue's NAPI if the queue has a
provider, so the driver tries to refill.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
Documentation/networking/netmem.rst | 7 ++-
include/net/page_pool/memory_provider.h | 11 +++++
include/net/page_pool/types.h | 7 +--
net/core/dev.c | 4 +-
net/core/netdev_rx_queue.c | 57 ++++++++++++++++++++-----
net/ethtool/rings.c | 18 +++-----
net/xdp/xsk_buff_pool.c | 3 +-
7 files changed, 78 insertions(+), 29 deletions(-)
diff --git a/Documentation/networking/netmem.rst b/Documentation/networking/netmem.rst
index 217869d1108d..aacd185843ac 100644
--- a/Documentation/networking/netmem.rst
+++ b/Documentation/networking/netmem.rst
@@ -48,8 +48,11 @@ Driver RX Requirements
- PP_FLAG_DMA_SYNC_DEV: netmem dma addr is not necessarily dma-syncable
by the driver. The driver must delegate the dma syncing to the page_pool,
which knows when dma-syncing is (or is not) appropriate.
- - PP_FLAG_ALLOW_UNREADABLE_NETMEM. The driver must specify this flag iff
- tcp-data-split is enabled.
+ - PP_FLAG_ALLOW_UNREADABLE_NETMEM. This opts the page pool into the memory
+ provider configured on its RX queue. For a provider without
+ MP_CAP_READABLE, the driver must specify it only for a payload pool
+ with tcp-data-split enabled. A readable provider may also back a
+ regular or header pool.
5. The driver must not assume the netmem is readable and/or backed by pages.
The netmem returned by the page_pool may be unreadable, in which case
diff --git a/include/net/page_pool/memory_provider.h b/include/net/page_pool/memory_provider.h
index 255ce4cfd975..d18ec079ffe2 100644
--- a/include/net/page_pool/memory_provider.h
+++ b/include/net/page_pool/memory_provider.h
@@ -9,6 +9,15 @@ struct netdev_rx_queue;
struct netlink_ext_ack;
struct sk_buff;
+/**
+ * enum mp_caps - memory provider capabilities
+ * @MP_CAP_READABLE: The CPU can access provider buffers. They may back
+ * header and regular page pools, and XDP programs may run on them.
+ */
+enum mp_caps {
+ MP_CAP_READABLE = BIT(0),
+};
+
struct memory_provider_ops {
netmem_ref (*alloc_netmems)(struct page_pool *pool, gfp_t gfp);
bool (*release_netmem)(struct page_pool *pool, netmem_ref netmem);
@@ -17,11 +26,13 @@ struct memory_provider_ops {
int (*nl_fill)(void *mp_priv, struct sk_buff *rsp,
struct netdev_rx_queue *rxq);
void (*uninstall)(void *mp_priv, struct netdev_rx_queue *rxq);
+ u32 caps;
};
bool net_mp_niov_set_dma_addr(struct net_iov *niov, dma_addr_t addr);
void net_mp_niov_set_page_pool(struct page_pool *pool, struct net_iov *niov);
void net_mp_niov_clear_page_pool(struct net_iov *niov);
+bool netif_mp_lacks_cap(struct net_device *dev, u32 cap);
int netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
const struct pp_memory_provider_params *p,
diff --git a/include/net/page_pool/types.h b/include/net/page_pool/types.h
index 03da138722f5..a96376613dda 100644
--- a/include/net/page_pool/types.h
+++ b/include/net/page_pool/types.h
@@ -22,9 +22,10 @@
*/
#define PP_FLAG_SYSTEM_POOL BIT(2) /* Global system page_pool */
-/* Allow unreadable (net_iov backed) netmem in this page_pool. Drivers setting
- * this must be able to support unreadable netmem, where netmem_address() would
- * return NULL. This flag should not be set for header page_pools.
+/* Allow memory-provider (net_iov backed) netmem in this page_pool. Drivers
+ * setting this must honor the provider's capabilities. Unless the provider has
+ * MP_CAP_READABLE, netmem_address() returns NULL, and the flag must not be set
+ * for a header page_pool.
*
* If the driver sets PP_FLAG_ALLOW_UNREADABLE_NETMEM, it should also set
* page_pool_params.slow.queue_idx.
diff --git a/net/core/dev.c b/net/core/dev.c
index 5ac08da8b9d7..12b532c09a9c 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -10423,7 +10423,7 @@ int netif_xdp_propagate(struct net_device *dev, struct netdev_bpf *bpf)
return -EBUSY;
}
- if (dev_get_min_mp_channel_count(dev)) {
+ if (netif_mp_lacks_cap(dev, MP_CAP_READABLE)) {
NL_SET_ERR_MSG(bpf->extack, "unable to propagate XDP to device using memory provider");
return -EBUSY;
}
@@ -10499,7 +10499,7 @@ static int dev_xdp_install(struct net_device *dev, enum bpf_xdp_mode mode,
return -EBUSY;
}
- if (dev_get_min_mp_channel_count(dev)) {
+ if (netif_mp_lacks_cap(dev, MP_CAP_READABLE)) {
NL_SET_ERR_MSG(extack, "unable to install XDP to device using memory provider");
return -EBUSY;
}
diff --git a/net/core/netdev_rx_queue.c b/net/core/netdev_rx_queue.c
index 00a7011eb4d5..610e31d1a717 100644
--- a/net/core/netdev_rx_queue.c
+++ b/net/core/netdev_rx_queue.c
@@ -82,8 +82,12 @@ __netif_get_rx_queue_lease(struct net_device **dev, unsigned int *rxq_idx,
/* See also page_pool_is_unreadable() */
bool netif_rxq_has_unreadable_mp(struct net_device *dev, unsigned int rxq_idx)
{
- if (rxq_idx < dev->real_num_rx_queues)
- return __netif_get_rx_queue(dev, rxq_idx)->mp_params.mp_ops;
+ const struct memory_provider_ops *ops;
+
+ if (rxq_idx < dev->real_num_rx_queues) {
+ ops = __netif_get_rx_queue(dev, rxq_idx)->mp_params.mp_ops;
+ return ops && !(ops->caps & MP_CAP_READABLE);
+ }
return false;
}
EXPORT_SYMBOL(netif_rxq_has_unreadable_mp);
@@ -95,6 +99,32 @@ bool netif_rxq_has_mp(struct net_device *dev, unsigned int rxq_idx)
return false;
}
+/* Return true if an installed memory provider lacks @cap. */
+bool netif_mp_lacks_cap(struct net_device *dev, u32 cap)
+{
+ const struct memory_provider_ops *ops;
+ int i;
+
+ netdev_assert_locked_ops_compat(dev);
+
+ for (i = dev->real_num_rx_queues - 1; i >= 0; i--) {
+ ops = dev->_rx[i].mp_params.mp_ops;
+ if (ops && (ops->caps & cap) != cap)
+ return true;
+ }
+
+ return false;
+}
+
+/* Provider memory may become available while the queue is replaced,
+ * without a device event to report it.
+ */
+static void netdev_rx_queue_kick(struct netdev_rx_queue *rxq)
+{
+ if (rxq->mp_params.mp_ops && rxq->napi)
+ napi_schedule(rxq->napi);
+}
+
static int netdev_rx_queue_reconfig(struct net_device *dev,
unsigned int rxq_idx,
struct netdev_queue_config *qcfg_old,
@@ -142,6 +172,8 @@ static int netdev_rx_queue_reconfig(struct net_device *dev,
}
qops->ndo_queue_mem_free(dev, old_mem);
+ if (netif_running(dev))
+ netdev_rx_queue_kick(rxq);
kvfree(old_mem);
kvfree(new_mem);
@@ -165,6 +197,11 @@ static int netdev_rx_queue_reconfig(struct net_device *dev,
err_free_new_queue_mem:
qops->ndo_queue_mem_free(dev, new_mem);
+ /* Freeing the replacement can hand provider memory back to the old
+ * queue, which still runs unless the rollback failed.
+ */
+ if (err != -ENETDOWN && netif_running(dev))
+ netdev_rx_queue_kick(rxq);
err_free_old_mem:
kvfree(old_mem);
@@ -189,6 +226,7 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
struct netlink_ext_ack *extack)
{
const struct netdev_queue_mgmt_ops *qops = dev->queue_mgmt_ops;
+ u32 caps = p->mp_ops->caps;
struct netdev_queue_config qcfg[2];
struct netdev_rx_queue *rxq;
int ret;
@@ -196,15 +234,13 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
if (!qops)
return -EOPNOTSUPP;
- if (dev->cfg->hds_config != ETHTOOL_TCP_DATA_SPLIT_ENABLED) {
- NL_SET_ERR_MSG(extack, "tcp-data-split is disabled");
+ if ((dev->cfg->hds_config != ETHTOOL_TCP_DATA_SPLIT_ENABLED ||
+ dev->cfg->hds_thresh) && !(caps & MP_CAP_READABLE)) {
+ NL_SET_ERR_MSG(extack,
+ "memory provider does not support current tcp-data-split configuration");
return -EINVAL;
}
- if (dev->cfg->hds_thresh) {
- NL_SET_ERR_MSG(extack, "hds-thresh is not zero");
- return -EINVAL;
- }
- if (dev_xdp_prog_count(dev)) {
+ if (dev_xdp_prog_count(dev) && !(caps & MP_CAP_READABLE)) {
NL_SET_ERR_MSG(extack, "unable to custom memory provider to device with XDP program attached");
return -EEXIST;
}
@@ -219,7 +255,8 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
return -EEXIST;
}
#ifdef CONFIG_XDP_SOCKETS
- if (rxq->pool) {
+ /* An AF_XDP pool owns the queue unless it is this provider. */
+ if (rxq->pool && rxq->pool != p->mp_priv) {
NL_SET_ERR_MSG(extack, "designated queue already in use by AF_XDP");
return -EBUSY;
}
diff --git a/net/ethtool/rings.c b/net/ethtool/rings.c
index e3810c0320e3..789d0e328b77 100644
--- a/net/ethtool/rings.c
+++ b/net/ethtool/rings.c
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: GPL-2.0-only
#include <net/netdev_queues.h>
+#include <net/page_pool/memory_provider.h>
#include "common.h"
#include "netlink.h"
@@ -258,17 +259,12 @@ ethnl_set_rings(struct ethnl_req_info *req_info, struct genl_info *info)
return -EINVAL;
}
- if (dev_get_min_mp_channel_count(dev)) {
- if (kernel_ringparam.tcp_data_split !=
- ETHTOOL_TCP_DATA_SPLIT_ENABLED) {
- NL_SET_ERR_MSG(info->extack,
- "can't disable tcp-data-split while device has memory provider enabled");
- return -EINVAL;
- } else if (kernel_ringparam.hds_thresh) {
- NL_SET_ERR_MSG(info->extack,
- "can't set non-zero hds_thresh while device is memory provider enabled");
- return -EINVAL;
- }
+ if ((kernel_ringparam.tcp_data_split !=
+ ETHTOOL_TCP_DATA_SPLIT_ENABLED || kernel_ringparam.hds_thresh) &&
+ netif_mp_lacks_cap(dev, MP_CAP_READABLE)) {
+ NL_SET_ERR_MSG(info->extack,
+ "memory provider does not support requested tcp-data-split configuration");
+ return -EINVAL;
}
/* ensure new ring parameters are within limits */
diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
index 9d2d94f1fb75..f01e1de7360e 100644
--- a/net/xdp/xsk_buff_pool.c
+++ b/net/xdp/xsk_buff_pool.c
@@ -2,6 +2,7 @@
#include <linux/netdevice.h>
#include <net/netdev_lock.h>
+#include <net/page_pool/memory_provider.h>
#include <net/xsk_buff_pool.h>
#include <net/xdp_sock.h>
#include <net/xdp_sock_drv.h>
@@ -235,7 +236,7 @@ int xp_assign_dev(struct xsk_buff_pool *pool,
goto err_unreg_pool;
}
- if (dev_get_min_mp_channel_count(netdev)) {
+ if (netif_mp_lacks_cap(netdev, MP_CAP_READABLE)) {
err = -EBUSY;
goto err_unreg_pool;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [RFC net-next 03/15] net: Add memory provider capabilities
2026-10-02 19:00 ` [RFC net-next 03/15] net: Add memory provider capabilities Björn Töpel
@ 2026-10-03 4:13 ` Mina Almasry
0 siblings, 0 replies; 17+ messages in thread
From: Mina Almasry @ 2026-10-03 4:13 UTC (permalink / raw)
To: Björn Töpel, Luigi Rizzo
Cc: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring, Mike Marciniszyn (Meta),
Weiming Shi, Nikolay Aleksandrov, David Wei, Alexander Lobakin,
linux-doc, linux-kernel
On Fri, Oct 2, 2026 at 12:01 PM Björn Töpel <bjorn@kernel.org> wrote:
>
> The current memory providers, devmem and io_uring zero-copy RX, give
> memory that the CPU cannot read. The core allows them only on queues
> with header-data split on, a zero header-split threshold, and no XDP
> program. AF_XDP needs a provider that the CPU can read and that
> works with XDP.
>
> Add a capability field to the provider operations, with
> MP_CAP_READABLE as the first flag. Apply the header-split and XDP
> rules only to providers without it. A queue with an AF_XDP buffer
> pool still rejects other providers, but it accepts the provider made
> from that same pool.
>
> Provider memory can become free while a queue is restarted, and no
> device event reports it. After a restart, and after a failed restart
> is rolled back, schedule the queue's NAPI if the queue has a
> provider, so the driver tries to refill.
>
> Signed-off-by: Björn Töpel <bjorn@kernel.org>
> ---
> Documentation/networking/netmem.rst | 7 ++-
> include/net/page_pool/memory_provider.h | 11 +++++
> include/net/page_pool/types.h | 7 +--
> net/core/dev.c | 4 +-
> net/core/netdev_rx_queue.c | 57 ++++++++++++++++++++-----
> net/ethtool/rings.c | 18 +++-----
> net/xdp/xsk_buff_pool.c | 3 +-
> 7 files changed, 78 insertions(+), 29 deletions(-)
>
> diff --git a/Documentation/networking/netmem.rst b/Documentation/networking/netmem.rst
> index 217869d1108d..aacd185843ac 100644
> --- a/Documentation/networking/netmem.rst
> +++ b/Documentation/networking/netmem.rst
> @@ -48,8 +48,11 @@ Driver RX Requirements
> - PP_FLAG_DMA_SYNC_DEV: netmem dma addr is not necessarily dma-syncable
> by the driver. The driver must delegate the dma syncing to the page_pool,
> which knows when dma-syncing is (or is not) appropriate.
> - - PP_FLAG_ALLOW_UNREADABLE_NETMEM. The driver must specify this flag iff
> - tcp-data-split is enabled.
> + - PP_FLAG_ALLOW_UNREADABLE_NETMEM. This opts the page pool into the memory
> + provider configured on its RX queue. For a provider without
> + MP_CAP_READABLE, the driver must specify it only for a payload pool
> + with tcp-data-split enabled. A readable provider may also back a
> + regular or header pool.
I kinda don't like this doc update. This flag was never meant to say
'I support memory providers'. It's just that the existing memory
providers are all unreadable so support-unreadable (headersplit) ==
supports-memory-providers.
The code should be updated such that if the memory provider is
!MP_CAP_READABLE and driver doesn't support
PP_FLAG_ALLOW_UNReADABLE_NETMEM, fail.
If another limitation applies to the new provider, add it as a
separate flag. The doc is accurate as-is I think.
>
> 5. The driver must not assume the netmem is readable and/or backed by pages.
> The netmem returned by the page_pool may be unreadable, in which case
> diff --git a/include/net/page_pool/memory_provider.h b/include/net/page_pool/memory_provider.h
> index 255ce4cfd975..d18ec079ffe2 100644
> --- a/include/net/page_pool/memory_provider.h
> +++ b/include/net/page_pool/memory_provider.h
> @@ -9,6 +9,15 @@ struct netdev_rx_queue;
> struct netlink_ext_ack;
> struct sk_buff;
>
> +/**
> + * enum mp_caps - memory provider capabilities
> + * @MP_CAP_READABLE: The CPU can access provider buffers. They may back
> + * header and regular page pools, and XDP programs may run on them.
> + */
> +enum mp_caps {
> + MP_CAP_READABLE = BIT(0),
> +};
> +
> struct memory_provider_ops {
> netmem_ref (*alloc_netmems)(struct page_pool *pool, gfp_t gfp);
> bool (*release_netmem)(struct page_pool *pool, netmem_ref netmem);
> @@ -17,11 +26,13 @@ struct memory_provider_ops {
> int (*nl_fill)(void *mp_priv, struct sk_buff *rsp,
> struct netdev_rx_queue *rxq);
> void (*uninstall)(void *mp_priv, struct netdev_rx_queue *rxq);
> + u32 caps;
> };
>
> bool net_mp_niov_set_dma_addr(struct net_iov *niov, dma_addr_t addr);
> void net_mp_niov_set_page_pool(struct page_pool *pool, struct net_iov *niov);
> void net_mp_niov_clear_page_pool(struct net_iov *niov);
> +bool netif_mp_lacks_cap(struct net_device *dev, u32 cap);
>
> int netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
> const struct pp_memory_provider_params *p,
> diff --git a/include/net/page_pool/types.h b/include/net/page_pool/types.h
> index 03da138722f5..a96376613dda 100644
> --- a/include/net/page_pool/types.h
> +++ b/include/net/page_pool/types.h
> @@ -22,9 +22,10 @@
> */
> #define PP_FLAG_SYSTEM_POOL BIT(2) /* Global system page_pool */
>
> -/* Allow unreadable (net_iov backed) netmem in this page_pool. Drivers setting
> - * this must be able to support unreadable netmem, where netmem_address() would
> - * return NULL. This flag should not be set for header page_pools.
> +/* Allow memory-provider (net_iov backed) netmem in this page_pool. Drivers
Comment is slightly inaccurate. It implies memory-providers can
only-even be net_iov backed. There is nothing wrong with a memory
provider returning page-backed-netmems. Mke it something like:
Allow unreadable memory-providers in this page_pool. Drivers that
support this must support unreadable netmem...
Have your pet LLM please go over the code for any instances in the
code/comments where we assumed net_iov == unreadable and fix those
with checks.
Here are the places my pet LLM thinks you missed:
### Handled by the Series
• netmem_address(): Returns net_iov_address() instead of NULL for
readable net_iovs.
• page_pool_is_unreadable() & netif_rxq_has_unreadable_mp(): Check
!(ops->caps & MP_CAP_READABLE).
• xdp_buff_add_frag(): Only sets XDP_FLAGS_FRAGS_UNREADABLE when
!net_iov_is_readable(niov).
### Missed by the Series (Feedback to Add)
1. xdp_build_skb_from_buff() (page head + readable net_iov frags):
Only checks xdp_buff_has_netmem(xdp) (head buffer), and
xdp_buff_add_frag() no longer sets XDP_FLAGS_FRAGS_UNREADABLE for
readable net_iovs.
A packet with a page-backed head and NET_IOV_XSK frags won't be
copied and will leak NET_IOV_XSK into an skb with skb->unreadable = 0.
2. skb_frag_address() & skb_frag_address_safe(): Still return NULL
if !skb_frag_page(frag). Patch 11 added a duplicate xdp_frag_address()
for 6 call sites, leaving other XDP frag readers (bpf_test_finish(),
driver multi-buffer XDP_TX like mlx5e_xdp_mpwqe_add_dseg()) broken.
skb_frag_address() itself should use netmem_address().
3. __skb_fill_netmem_desc() & skb_dump(): Still treat all net_iovs
as unreadable (skb->unreadable = true). Also, drivers that build skbs
via skb_add_rx_frag_netmem() when !xdp_prog will attach NET_IOV_XSK
directly to skbs (which __get_netmem()` / `__put_netmem() don't refcount).
4. bnxt & bnge need_head_pool: Set rxr->need_head_pool =
page_pool_is_unreadable(pool) to decide whether head_pool needs a
separate struct page pool, then call page-only page_pool_alloc_frag()
and
page_pool_free_va() on head_pool. Returning false from
page_pool_is_unreadable() breaks them.
5. Drivers checking netmem_is_net_iov() as "unreadable":
• gve_rx_dqo.c:892, 953: Uses netmem_is_net_iov() to drop
unsplit headers and skip copybreak; should check !netmem_address().
• en_rx.c:2282-2286: Uses netmem_is_net_iov() to drop unsplit
headers; should check !netmem_address().
• en_main.c:5694-5705: Uses netif_rxq_has_unreadable_mp() to
validate custom page size; should be netif_rxq_has_mp().
• idpf_txrx.c:3470-3479: Uses netmem_is_net_iov() and
__netmem_to_page() in idpf_rx_hsplit_wa(); should use
netmem_address().
6. Stale comments/docs:
• netmem.h:73-83: Comment on struct net_iov still says it is
only for non-struct page memory and fixed PAGE_SIZE chunks.
• netmem.rst:27 & types.h:32: Still state tcp-data-split is
unconditionally required and keep the PP_FLAG_ALLOW_UNREADABLE_NETMEM
name for readable providers.
> + * setting this must honor the provider's capabilities. Unless the provider has
> + * MP_CAP_READABLE, netmem_address() returns NULL, and the flag must not be set
> + * for a header page_pool.
> *
That bit about header page_pool is good to add anyway.
> * If the driver sets PP_FLAG_ALLOW_UNREADABLE_NETMEM, it should also set
> * page_pool_params.slow.queue_idx.
> diff --git a/net/core/dev.c b/net/core/dev.c
> index 5ac08da8b9d7..12b532c09a9c 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -10423,7 +10423,7 @@ int netif_xdp_propagate(struct net_device *dev, struct netdev_bpf *bpf)
> return -EBUSY;
> }
>
> - if (dev_get_min_mp_channel_count(dev)) {
> + if (netif_mp_lacks_cap(dev, MP_CAP_READABLE)) {
Could use a name easier to read?
if (!netif_all_mps_have_cap())
> NL_SET_ERR_MSG(bpf->extack, "unable to propagate XDP to device using memory provider");
...using unreadable memory provider");
> return -EBUSY;
> }
> @@ -10499,7 +10499,7 @@ static int dev_xdp_install(struct net_device *dev, enum bpf_xdp_mode mode,
> return -EBUSY;
> }
>
> - if (dev_get_min_mp_channel_count(dev)) {
> + if (netif_mp_lacks_cap(dev, MP_CAP_READABLE)) {
> NL_SET_ERR_MSG(extack, "unable to install XDP to device using memory provider");
Same.
> return -EBUSY;
> }
> diff --git a/net/core/netdev_rx_queue.c b/net/core/netdev_rx_queue.c
> index 00a7011eb4d5..610e31d1a717 100644
> --- a/net/core/netdev_rx_queue.c
> +++ b/net/core/netdev_rx_queue.c
> @@ -82,8 +82,12 @@ __netif_get_rx_queue_lease(struct net_device **dev, unsigned int *rxq_idx,
> /* See also page_pool_is_unreadable() */
> bool netif_rxq_has_unreadable_mp(struct net_device *dev, unsigned int rxq_idx)
> {
> - if (rxq_idx < dev->real_num_rx_queues)
> - return __netif_get_rx_queue(dev, rxq_idx)->mp_params.mp_ops;
> + const struct memory_provider_ops *ops;
> +
> + if (rxq_idx < dev->real_num_rx_queues) {
> + ops = __netif_get_rx_queue(dev, rxq_idx)->mp_params.mp_ops;
> + return ops && !(ops->caps & MP_CAP_READABLE);
> + }
handle error case inside if() block please.
> return false;
> }
> EXPORT_SYMBOL(netif_rxq_has_unreadable_mp);
> @@ -95,6 +99,32 @@ bool netif_rxq_has_mp(struct net_device *dev, unsigned int rxq_idx)
> return false;
> }
>
> +/* Return true if an installed memory provider lacks @cap. */
> +bool netif_mp_lacks_cap(struct net_device *dev, u32 cap)
> +{
> + const struct memory_provider_ops *ops;
> + int i;
> +
> + netdev_assert_locked_ops_compat(dev);
> +
> + for (i = dev->real_num_rx_queues - 1; i >= 0; i--) {
> + ops = dev->_rx[i].mp_params.mp_ops;
> + if (ops && (ops->caps & cap) != cap)
> + return true;
> + }
> +
> + return false;
> +}
> +
> +/* Provider memory may become available while the queue is replaced,
> + * without a device event to report it.
> + */
> +static void netdev_rx_queue_kick(struct netdev_rx_queue *rxq)
> +{
> + if (rxq->mp_params.mp_ops && rxq->napi)
> + napi_schedule(rxq->napi);
> +}
> +
Honestly maybe open code this at the call site. It's confusing to have
a hepler that doesn't mention mp only do something if the rx queue has
an mp configured. Or maybe netdev_mp_rx_queue_kick().
> static int netdev_rx_queue_reconfig(struct net_device *dev,
> unsigned int rxq_idx,
> struct netdev_queue_config *qcfg_old,
> @@ -142,6 +172,8 @@ static int netdev_rx_queue_reconfig(struct net_device *dev,
> }
>
> qops->ndo_queue_mem_free(dev, old_mem);
> + if (netif_running(dev))
> + netdev_rx_queue_kick(rxq);
>
> kvfree(old_mem);
> kvfree(new_mem);
> @@ -165,6 +197,11 @@ static int netdev_rx_queue_reconfig(struct net_device *dev,
>
> err_free_new_queue_mem:
> qops->ndo_queue_mem_free(dev, new_mem);
> + /* Freeing the replacement can hand provider memory back to the old
> + * queue, which still runs unless the rollback failed.
> + */
> + if (err != -ENETDOWN && netif_running(dev))
> + netdev_rx_queue_kick(rxq);
These kicks feel extremely error prone and hard to point where in the
code we missed a necessary kick. Try to find something better to do.
>
> err_free_old_mem:
> kvfree(old_mem);
> @@ -189,6 +226,7 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
> struct netlink_ext_ack *extack)
> {
> const struct netdev_queue_mgmt_ops *qops = dev->queue_mgmt_ops;
> + u32 caps = p->mp_ops->caps;
> struct netdev_queue_config qcfg[2];
> struct netdev_rx_queue *rxq;
> int ret;
> @@ -196,15 +234,13 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
> if (!qops)
> return -EOPNOTSUPP;
>
> - if (dev->cfg->hds_config != ETHTOOL_TCP_DATA_SPLIT_ENABLED) {
> - NL_SET_ERR_MSG(extack, "tcp-data-split is disabled");
> + if ((dev->cfg->hds_config != ETHTOOL_TCP_DATA_SPLIT_ENABLED ||
> + dev->cfg->hds_thresh) && !(caps & MP_CAP_READABLE)) {
nit: This could use a helper mp_caps_readable() or something.
> + NL_SET_ERR_MSG(extack,
> + "memory provider does not support current tcp-data-split configuration");
> return -EINVAL;
> }
> - if (dev->cfg->hds_thresh) {
> - NL_SET_ERR_MSG(extack, "hds-thresh is not zero");
> - return -EINVAL;
> - }
> - if (dev_xdp_prog_count(dev)) {
> + if (dev_xdp_prog_count(dev) && !(caps & MP_CAP_READABLE)) {
> NL_SET_ERR_MSG(extack, "unable to custom memory provider to device with XDP program attached");
unable to use unreadable memory provider with XDP program attached maybe.
> return -EEXIST;
> }
> @@ -219,7 +255,8 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
> return -EEXIST;
> }
> #ifdef CONFIG_XDP_SOCKETS
> - if (rxq->pool) {
> + /* An AF_XDP pool owns the queue unless it is this provider. */
> + if (rxq->pool && rxq->pool != p->mp_priv) {
> NL_SET_ERR_MSG(extack, "designated queue already in use by AF_XDP");
> return -EBUSY;
> }
> diff --git a/net/ethtool/rings.c b/net/ethtool/rings.c
> index e3810c0320e3..789d0e328b77 100644
> --- a/net/ethtool/rings.c
> +++ b/net/ethtool/rings.c
> @@ -1,6 +1,7 @@
> // SPDX-License-Identifier: GPL-2.0-only
>
> #include <net/netdev_queues.h>
> +#include <net/page_pool/memory_provider.h>
>
> #include "common.h"
> #include "netlink.h"
> @@ -258,17 +259,12 @@ ethnl_set_rings(struct ethnl_req_info *req_info, struct genl_info *info)
> return -EINVAL;
> }
>
> - if (dev_get_min_mp_channel_count(dev)) {
> - if (kernel_ringparam.tcp_data_split !=
> - ETHTOOL_TCP_DATA_SPLIT_ENABLED) {
> - NL_SET_ERR_MSG(info->extack,
> - "can't disable tcp-data-split while device has memory provider enabled");
> - return -EINVAL;
> - } else if (kernel_ringparam.hds_thresh) {
> - NL_SET_ERR_MSG(info->extack,
> - "can't set non-zero hds_thresh while device is memory provider enabled");
> - return -EINVAL;
> - }
> + if ((kernel_ringparam.tcp_data_split !=
> + ETHTOOL_TCP_DATA_SPLIT_ENABLED || kernel_ringparam.hds_thresh) &&
> + netif_mp_lacks_cap(dev, MP_CAP_READABLE)) {
> + NL_SET_ERR_MSG(info->extack,
> + "memory provider does not support requested tcp-data-split configuration");
> + return -EINVAL;
> }
>
> /* ensure new ring parameters are within limits */
> diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
> index 9d2d94f1fb75..f01e1de7360e 100644
> --- a/net/xdp/xsk_buff_pool.c
> +++ b/net/xdp/xsk_buff_pool.c
> @@ -2,6 +2,7 @@
>
> #include <linux/netdevice.h>
> #include <net/netdev_lock.h>
> +#include <net/page_pool/memory_provider.h>
> #include <net/xsk_buff_pool.h>
> #include <net/xdp_sock.h>
> #include <net/xdp_sock_drv.h>
> @@ -235,7 +236,7 @@ int xp_assign_dev(struct xsk_buff_pool *pool,
> goto err_unreg_pool;
> }
>
> - if (dev_get_min_mp_channel_count(netdev)) {
> + if (netif_mp_lacks_cap(netdev, MP_CAP_READABLE)) {
> err = -EBUSY;
> goto err_unreg_pool;
> }
> --
> 2.55.0
>
--
Thanks,
Mina
^ permalink raw reply [flat|nested] 17+ messages in thread
* [RFC net-next 04/15] net: Let memory providers set RX buffer headroom
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (2 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 03/15] net: Add memory provider capabilities Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 05/15] page_pool: Extend memory provider operations Björn Töpel
` (10 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
AF_XDP lets userspace choose a headroom in front of the packet data
in each UMEM chunk, and the kernel adds XDP_PACKET_HEADROOM to it. A
driver whose buffers come from an AF_XDP memory provider must put
packets after this headroom. The driver sees only the page pool, so
it has no generic way to learn the value, and nothing checks that
the driver supports it.
Add rx_headroom to the provider parameters and to the queue
configuration. It is the full headroom, including
XDP_PACKET_HEADROOM; zero means the driver default. Add
QCFG_RX_HEADROOM. A provider that sets a headroom can only be
installed on a queue that supports it; otherwise the install fails
with -EOPNOTSUPP. The driver checks the value with the rest of the
queue configuration.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/netdev_queues.h | 6 ++++++
include/net/page_pool/types.h | 1 +
net/core/netdev_config.c | 2 ++
net/core/netdev_rx_queue.c | 4 ++++
4 files changed, 13 insertions(+)
diff --git a/include/net/netdev_queues.h b/include/net/netdev_queues.h
index c5335e935b7d..f9eba63a5d78 100644
--- a/include/net/netdev_queues.h
+++ b/include/net/netdev_queues.h
@@ -45,12 +45,16 @@ struct netdev_config {
/**
* struct netdev_queue_config - rendered configuration for an RX queue
* @rx_page_size: Size of one RX page-pool allocation.
+ * @rx_headroom: Headroom before packet data in the first buffer,
+ * including XDP_PACKET_HEADROOM. Zero selects the
+ * driver default.
* @rx_ring_size: Configured size of the regular RX ring.
* @rx_mini_ring_size: Configured size of the RX mini ring.
* @rx_jumbo_ring_size: Configured size of the RX jumbo ring.
*/
struct netdev_queue_config {
u32 rx_page_size;
+ u32 rx_headroom;
u32 rx_ring_size;
u32 rx_mini_ring_size;
u32 rx_jumbo_ring_size;
@@ -156,6 +160,8 @@ void netdev_stat_queue_sum(struct net_device *netdev,
enum {
/* The queue checks and honours the page size qcfg parameter */
QCFG_RX_PAGE_SIZE = 0x1,
+ /* The queue checks and honours the headroom qcfg parameter */
+ QCFG_RX_HEADROOM = 0x2,
};
/**
diff --git a/include/net/page_pool/types.h b/include/net/page_pool/types.h
index a96376613dda..a673fa35febb 100644
--- a/include/net/page_pool/types.h
+++ b/include/net/page_pool/types.h
@@ -172,6 +172,7 @@ struct pp_memory_provider_params {
void *mp_priv;
const struct memory_provider_ops *mp_ops;
u32 rx_page_size;
+ u32 rx_headroom;
};
struct page_pool {
diff --git a/net/core/netdev_config.c b/net/core/netdev_config.c
index 1975de42a60d..97e33764737b 100644
--- a/net/core/netdev_config.c
+++ b/net/core/netdev_config.c
@@ -88,6 +88,8 @@ static int __netdev_queue_config(struct net_device *dev, int rxq_idx,
mpp = &__netif_get_rx_queue(dev, rxq_idx)->mp_params;
if (mpp->rx_page_size)
qcfg->rx_page_size = mpp->rx_page_size;
+ if (mpp->rx_headroom)
+ qcfg->rx_headroom = mpp->rx_headroom;
err = validate_cb(dev, qcfg, extack);
if (err)
return err;
diff --git a/net/core/netdev_rx_queue.c b/net/core/netdev_rx_queue.c
index 610e31d1a717..476289000e78 100644
--- a/net/core/netdev_rx_queue.c
+++ b/net/core/netdev_rx_queue.c
@@ -248,6 +248,10 @@ static int __netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
NL_SET_ERR_MSG(extack, "device does not support: rx_page_size");
return -EOPNOTSUPP;
}
+ if (p->rx_headroom && !(qops->supported_params & QCFG_RX_HEADROOM)) {
+ NL_SET_ERR_MSG(extack, "device does not support: rx_headroom");
+ return -EOPNOTSUPP;
+ }
rxq = __netif_get_rx_queue(dev, rxq_idx);
if (rxq->mp_params.mp_ops) {
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 05/15] page_pool: Extend memory provider operations
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (3 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 04/15] net: Let memory providers set RX buffer headroom Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 06/15] xdp: Track non-page netmem in receive buffers Björn Töpel
` (9 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
The current memory providers give memory that the CPU cannot read,
and they can always refill. An AF_XDP provider is different. The CPU
can read its memory, it has only the buffers that userspace puts in
the FILL ring, and its buffers go back to userspace instead of being
freed. It needs four things from page_pool:
- A CPU address for each net_iov. A provider can set the address of
the first net_iov in its net_iov_area and the size each net_iov
covers. netmem_address() then computes the address without a call
into the provider; drivers call it for every packet. A
net_iov_area without an address stays unreadable.
page_pool_is_unreadable() is false for a readable provider.
- A way to say that the refill is not done. The new refill_done
callback tells the driver whether it may stop refilling. A
provider that still has buffers returns false, and the driver
keeps NAPI scheduled.
- A way to give up a buffer without freeing it. Add a batched call
that ends page_pool ownership of buffers that leave through the
provider, for example to userspace. It clears the page_pool link,
so a later page pool can take the buffer. Add the matching batched
call that sets the link, and a DMA sync helper for drivers that
post a buffer again directly.
- No buffer splitting, because a split buffer has several owners.
Add MP_CAP_FRAG. page_pool refuses fragment allocation from a
provider without it.
devmem and io_uring set MP_CAP_FRAG and keep their fragment
behaviour. Their refill paths now use the batched call that sets the
link.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/netmem.h | 31 ++++++++++-
include/net/page_pool/helpers.h | 62 +++++++++++++++++++++-
include/net/page_pool/memory_provider.h | 28 +++++-----
io_uring/zcrx.c | 4 +-
net/core/devmem.c | 7 ++-
net/core/page_pool.c | 68 ++++++++++++++++++++-----
6 files changed, 165 insertions(+), 35 deletions(-)
diff --git a/include/net/netmem.h b/include/net/netmem.h
index cc97611632dc..e6dff0b01581 100644
--- a/include/net/netmem.h
+++ b/include/net/netmem.h
@@ -105,6 +105,12 @@ struct net_iov_area {
/* Offset into the dma-buf where this chunk starts. */
unsigned long base_virtual;
+
+ /* CPU address of the first net_iov's memory, or NULL when the CPU
+ * cannot read the area. Each net_iov covers 1 << @niov_shift bytes.
+ */
+ void *vaddr;
+ u8 niov_shift;
};
static inline struct net_iov_area *net_iov_owner(const struct net_iov *niov)
@@ -117,6 +123,23 @@ static inline unsigned int net_iov_idx(const struct net_iov *niov)
return niov - net_iov_owner(niov)->niovs;
}
+static inline bool net_iov_is_readable(const struct net_iov *niov)
+{
+ return net_iov_owner(niov)->vaddr;
+}
+
+static inline void *net_iov_address(const struct net_iov *niov)
+{
+ const struct net_iov_area *area = net_iov_owner(niov);
+ unsigned long off;
+
+ if (!area->vaddr)
+ return NULL;
+
+ off = (unsigned long)net_iov_idx(niov) << area->niov_shift;
+ return area->vaddr + off;
+}
+
/* Initialize a niov: stamp the owning area, the memory provider type.
*/
static inline void net_iov_init(struct net_iov *niov,
@@ -335,10 +358,16 @@ static inline void *__netmem_address(netmem_ref netmem)
return page_address(__netmem_to_page(netmem));
}
+/**
+ * netmem_address - get pointer to the memory backing @netmem
+ * @netmem: netmem reference to get the pointer for
+ *
+ * Return: pointer to the memory, or NULL if the CPU cannot read @netmem.
+ */
static inline void *netmem_address(netmem_ref netmem)
{
if (netmem_is_net_iov(netmem))
- return NULL;
+ return net_iov_address(netmem_to_net_iov(netmem));
return __netmem_address(netmem);
}
diff --git a/include/net/page_pool/helpers.h b/include/net/page_pool/helpers.h
index cd021832c3fa..7a63671aaceb 100644
--- a/include/net/page_pool/helpers.h
+++ b/include/net/page_pool/helpers.h
@@ -54,6 +54,7 @@
#include <linux/dma-mapping.h>
+#include <net/page_pool/memory_provider.h>
#include <net/page_pool/types.h>
#include <net/net_debug.h>
#include <net/netmem.h>
@@ -427,6 +428,26 @@ static inline dma_addr_t page_pool_get_dma_addr_netmem(netmem_ref netmem)
return netmem_dma_addr_decode(netmem_get_dma_addr(netmem));
}
+/**
+ * page_pool_refill_done - finish a provider-backed RX ring refill
+ * @pool: page pool used by the RX ring
+ * @full: all RX rings sharing this page pool were refilled
+ *
+ * Return: true if the caller may stop retrying the refill.
+ */
+static inline bool page_pool_refill_done(struct page_pool *pool, bool full)
+{
+ if (unlikely(pool->mp_ops && pool->mp_ops->refill_done))
+ return pool->mp_ops->refill_done(pool, full);
+
+ return true;
+}
+
+static inline bool page_pool_supports_frag(const struct page_pool *pool)
+{
+ return !pool->mp_ops || pool->mp_ops->caps & MP_CAP_FRAG;
+}
+
/**
* page_pool_get_dma_addr() - Retrieve the stored DMA address.
* @page: page allocated from a page pool
@@ -481,6 +502,45 @@ page_pool_dma_sync_netmem_for_cpu(const struct page_pool *pool,
offset, dma_sync_size);
}
+/**
+ * page_pool_dma_sync_netmem_for_device - sync netmem before giving it to HW
+ * @pool: page pool the netmem belongs to
+ * @netmem: netmem to sync
+ * @offset: offset from the pool's DMA sync start
+ * @dma_sync_size: size of the memory area to sync
+ *
+ * Use when a driver reposts a netmem directly instead of returning it through
+ * page_pool_put_netmem().
+ */
+static inline void
+page_pool_dma_sync_netmem_for_device(const struct page_pool *pool,
+ const netmem_ref netmem, u32 offset,
+ u32 dma_sync_size)
+{
+#if defined(CONFIG_HAS_DMA) && defined(CONFIG_DMA_NEED_SYNC)
+ dma_addr_t dma_addr;
+
+ if (!pool->dma_sync || !dma_dev_need_sync(pool->p.dev))
+ return;
+
+ rcu_read_lock();
+ /* Recheck under RCU to synchronize with page_pool_scrub(). */
+ if (pool->dma_sync) {
+ if (WARN_ON_ONCE(offset > pool->p.max_len))
+ goto out;
+
+ dma_addr = page_pool_get_dma_addr_netmem(netmem);
+ dma_sync_size = min(dma_sync_size, pool->p.max_len - offset);
+ dma_sync_single_range_for_device(pool->p.dev, dma_addr,
+ offset + pool->p.offset,
+ dma_sync_size,
+ pool->p.dma_dir);
+ }
+out:
+ rcu_read_unlock();
+#endif
+}
+
static inline void page_pool_get(struct page_pool *pool)
{
refcount_inc(&pool->user_cnt);
@@ -511,7 +571,7 @@ static inline void page_pool_nid_changed(struct page_pool *pool, int new_nid)
*/
static inline bool page_pool_is_unreadable(struct page_pool *pool)
{
- return !!pool->mp_ops;
+ return pool->mp_ops && !(pool->mp_ops->caps & MP_CAP_READABLE);
}
#endif /* _NET_PAGE_POOL_HELPERS_H */
diff --git a/include/net/page_pool/memory_provider.h b/include/net/page_pool/memory_provider.h
index d18ec079ffe2..61768f2592c9 100644
--- a/include/net/page_pool/memory_provider.h
+++ b/include/net/page_pool/memory_provider.h
@@ -13,14 +13,22 @@ struct sk_buff;
* enum mp_caps - memory provider capabilities
* @MP_CAP_READABLE: The CPU can access provider buffers. They may back
* header and regular page pools, and XDP programs may run on them.
+ * @MP_CAP_FRAG: Page pool fragments may split provider buffers. A
+ * provider without it hands out objects with a reference count of one.
*/
enum mp_caps {
MP_CAP_READABLE = BIT(0),
+ MP_CAP_FRAG = BIT(1),
};
struct memory_provider_ops {
netmem_ref (*alloc_netmems)(struct page_pool *pool, gfp_t gfp);
bool (*release_netmem)(struct page_pool *pool, netmem_ref netmem);
+ /* Called after RX refill. @full says that the caller reached its refill
+ * target. Return true when NAPI may complete, or false to keep it
+ * scheduled even without packet completions.
+ */
+ bool (*refill_done)(struct page_pool *pool, bool full);
int (*init)(struct page_pool *pool);
void (*destroy)(struct page_pool *pool);
int (*nl_fill)(void *mp_priv, struct sk_buff *rsp,
@@ -30,7 +38,10 @@ struct memory_provider_ops {
};
bool net_mp_niov_set_dma_addr(struct net_iov *niov, dma_addr_t addr);
-void net_mp_niov_set_page_pool(struct page_pool *pool, struct net_iov *niov);
+void net_mp_netmem_set_page_pool_bulk(struct page_pool *pool,
+ netmem_ref *netmems, u32 count);
+void net_mp_release_page_pool_bulk(struct page_pool *pool,
+ netmem_ref *netmems, u32 count);
void net_mp_niov_clear_page_pool(struct net_iov *niov);
bool netif_mp_lacks_cap(struct net_device *dev, u32 cap);
@@ -40,19 +51,4 @@ int netif_mp_open_rxq(struct net_device *dev, unsigned int rxq_idx,
void netif_mp_close_rxq(struct net_device *dev, unsigned int rxq_idx,
const struct pp_memory_provider_params *old_p);
-/**
- * net_mp_netmem_place_in_cache() - give a netmem to a page pool
- * @pool: the page pool to place the netmem into
- * @netmem: netmem to give
- *
- * Push an accounted netmem into the page pool's allocation cache. The caller
- * must ensure that there is space in the cache. It should only be called off
- * the mp_ops->alloc_netmems() path.
- */
-static inline void net_mp_netmem_place_in_cache(struct page_pool *pool,
- netmem_ref netmem)
-{
- pool->alloc.cache[pool->alloc.count++] = netmem;
-}
-
#endif
diff --git a/io_uring/zcrx.c b/io_uring/zcrx.c
index 86d580d4410d..ffffbccd517f 100644
--- a/io_uring/zcrx.c
+++ b/io_uring/zcrx.c
@@ -1315,10 +1315,11 @@ static unsigned io_zcrx_refill_slow(struct page_pool *pp, struct io_zcrx_ifq *if
continue;
}
- net_mp_niov_set_page_pool(pp, niov);
netmems[allocated] = net_iov_to_netmem(niov);
allocated++;
}
+ if (allocated)
+ net_mp_netmem_set_page_pool_bulk(pp, netmems, allocated);
return allocated;
}
@@ -1476,6 +1477,7 @@ static const struct memory_provider_ops io_uring_pp_zc_ops = {
.destroy = io_pp_zc_destroy,
.nl_fill = io_pp_nl_fill,
.uninstall = io_pp_uninstall,
+ .caps = MP_CAP_FRAG,
};
static unsigned zcrx_parse_rq(netmem_ref *netmem_array, unsigned nr,
diff --git a/net/core/devmem.c b/net/core/devmem.c
index 71c83730d73d..aeb804662668 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -459,7 +459,7 @@ netmem_ref mp_dmabuf_devmem_alloc_netmems(struct page_pool *pool, gfp_t gfp)
{
struct net_devmem_dmabuf_binding *binding = pool->mp_priv;
netmem_ref *netmems = pool->alloc.cache;
- unsigned int allocated, i;
+ unsigned int allocated;
if (WARN_ON_ONCE(pool->alloc.count))
return 0;
@@ -469,9 +469,7 @@ netmem_ref mp_dmabuf_devmem_alloc_netmems(struct page_pool *pool, gfp_t gfp)
if (unlikely(!allocated))
return 0;
- for (i = 0; i < allocated; i++)
- net_mp_niov_set_page_pool(pool,
- netmem_to_net_iov(netmems[i]));
+ net_mp_netmem_set_page_pool_bulk(pool, netmems, allocated);
/* Return the last one, the rest stay in the page_pool cache. */
allocated--;
@@ -540,4 +538,5 @@ static const struct memory_provider_ops dmabuf_devmem_ops = {
.release_netmem = mp_dmabuf_devmem_release_page,
.nl_fill = mp_dmabuf_devmem_nl_fill,
.uninstall = mp_dmabuf_devmem_uninstall,
+ .caps = MP_CAP_FRAG,
};
diff --git a/net/core/page_pool.c b/net/core/page_pool.c
index d7c88c0b67e5..e36ae6123adf 100644
--- a/net/core/page_pool.c
+++ b/net/core/page_pool.c
@@ -726,11 +726,9 @@ void page_pool_set_pp_info(struct page_pool *pool, netmem_ref netmem)
netmem_set_pp(netmem, pool);
netmem_or_pp_magic(netmem, PP_SIGNATURE);
- /* Ensuring all pages have been split into one fragment initially:
- * page_pool_set_pp_info() is only called once for every page when it
- * is allocated from the page allocator and page_pool_fragment_page()
- * is dirtying the same cache line as the page->pp_magic above, so
- * the overhead is negligible.
+ /* Ensure all netmem starts with one fragment. The refcount and
+ * page-pool state initialized above share a cache line, so the overhead
+ * is negligible when an object is first associated with the pool.
*/
page_pool_fragment_netmem(netmem, 1);
if (pool->has_init_callback)
@@ -1070,6 +1068,9 @@ netmem_ref page_pool_alloc_frag_netmem(struct page_pool *pool,
unsigned int max_size = PAGE_SIZE << pool->p.order;
netmem_ref netmem = pool->frag_page;
+ if (static_branch_unlikely(&page_pool_mem_providers) &&
+ WARN_ON_ONCE(!page_pool_supports_frag(pool)))
+ return 0;
if (WARN_ON(size > max_size))
return 0;
@@ -1340,17 +1341,60 @@ bool net_mp_niov_set_dma_addr(struct net_iov *niov, dma_addr_t addr)
return page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov), addr);
}
-/* Associate a niov with a page pool. Should follow with a matching
- * net_mp_niov_clear_page_pool()
+/* Associate a batch of niovs with a page pool. Each needs a matching
+ * net_mp_niov_clear_page_pool() or net_mp_release_page_pool_bulk().
*/
-void net_mp_niov_set_page_pool(struct page_pool *pool, struct net_iov *niov)
+void net_mp_netmem_set_page_pool_bulk(struct page_pool *pool,
+ netmem_ref *netmems, u32 count)
{
- netmem_ref netmem = net_iov_to_netmem(niov);
+ bool trace = trace_page_pool_state_hold_enabled();
+ netmem_ref netmem;
+ u32 i;
- page_pool_set_pp_info(pool, netmem);
+ for (i = 0; i < count; i++) {
+ netmem = netmems[i];
+ page_pool_set_pp_info(pool, netmem);
+ if (!trace)
+ continue;
- pool->pages_state_hold_cnt++;
- trace_page_pool_state_hold(pool, netmem, pool->pages_state_hold_cnt);
+ pool->pages_state_hold_cnt++;
+ trace_page_pool_state_hold(pool, netmem,
+ pool->pages_state_hold_cnt);
+ }
+ if (!trace)
+ pool->pages_state_hold_cnt += count;
+}
+
+/* Release page_pool ownership of netmems which a provider hands out of the
+ * page pool, for example to userspace. All netmems in the batch must belong
+ * to @pool. The pool may be freed once the last release is accounted.
+ */
+void net_mp_release_page_pool_bulk(struct page_pool *pool,
+ netmem_ref *netmems, u32 count)
+{
+ bool trace = trace_page_pool_state_release_enabled();
+ atomic_t *release_cnt;
+ netmem_ref netmem;
+ int released;
+ u32 i;
+
+ if (WARN_ON_ONCE(!pool || !count))
+ return;
+
+ release_cnt = &pool->pages_state_release_cnt;
+
+ for (i = 0; i < count; i++) {
+ netmem = netmems[i];
+ DEBUG_NET_WARN_ON_ONCE(netmem_get_pp(netmem) != pool);
+ page_pool_clear_pp_info(netmem);
+ if (!trace)
+ continue;
+
+ released = atomic_inc_return_relaxed(release_cnt);
+ trace_page_pool_state_release(pool, netmem, released);
+ }
+ if (!trace)
+ atomic_add(count, release_cnt);
}
/* Disassociate a niov from a page pool. Should only be used in the
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 06/15] xdp: Track non-page netmem in receive buffers
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (4 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 05/15] page_pool: Extend memory provider operations Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 07/15] xsk: Keep the DMA mapping in the buffer pool Björn Töpel
` (8 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
A page-pool memory provider can give net_iov buffers, and
virt_to_page() does not work for them. An xdp_buff holds only the
data pointer, so code that returns or converts a buffer cannot find
its net_iov.
Add NET_IOV_XSK, the net_iov type for AF_XDP. Add the internal flag
XDP_FLAGS_HAS_NETMEM. When it is set, the xdp_buff holds a netmem
reference in the union with the devmap TX queue, which only devmap
egress uses. The flag is cleared when flags are copied to an skb or
an xdp_frame. xdp_convert_buff_to_frame() sends such a buffer to the
zero-copy copy path.
A fragment from a readable net_iov area no longer marks the buffer's
fragments as unreadable. The skb layer still treats every net_iov as
unreadable, so such a buffer must be copied before it becomes an
skb. Later patches in this series do that.
skb_shared_info holds kernel pointers, so it cannot be in UMEM,
which userspace can write. A netmem xdp_buff points to a
skb_shared_info that the RX queue keeps in kernel memory instead.
xdp_init_buff_from_netmem() sets this up. Page-backed buffers still
use their tailroom.
On 64-bit, struct xdp_buff grows from 56 to 64 bytes. struct
xdp_buff_xsk stays at 128 bytes on x86-64. Where the largest
alignment is 8 bytes, as on s390, it grows from 120 to 128 bytes, so
a pool needs 8 more bytes per UMEM chunk. Only helpers that need the
netmem or the fragment info test the new flag. Page-backed receive
and return paths do not.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/netmem.h | 1 +
include/net/xdp.h | 56 ++++++++++++++++++++++++++++++++++++++++----
2 files changed, 52 insertions(+), 5 deletions(-)
diff --git a/include/net/netmem.h b/include/net/netmem.h
index e6dff0b01581..76e4fe878863 100644
--- a/include/net/netmem.h
+++ b/include/net/netmem.h
@@ -68,6 +68,7 @@ DECLARE_STATIC_KEY_FALSE(page_pool_mem_providers);
enum net_iov_type {
NET_IOV_DMABUF,
NET_IOV_IOURING,
+ NET_IOV_XSK,
};
/* A memory descriptor representing abstract networking I/O vectors,
diff --git a/include/net/xdp.h b/include/net/xdp.h
index 07231adfb5f8..80931490ca3a 100644
--- a/include/net/xdp.h
+++ b/include/net/xdp.h
@@ -11,6 +11,7 @@
#include <linux/netdevice.h>
#include <linux/skbuff.h> /* skb_shared_info */
+#include <net/netmem.h>
#include <net/page_pool/types.h>
/**
@@ -81,6 +82,8 @@ enum xdp_buff_flags {
* XDP program is not attached.
*/
XDP_FLAGS_FRAGS_UNREADABLE = BIT(2),
+ /* The txq/netmem union contains a receive netmem reference. */
+ XDP_FLAGS_HAS_NETMEM = BIT(3),
};
struct xdp_buff {
@@ -89,7 +92,12 @@ struct xdp_buff {
void *data_meta;
void *data_hard_start;
struct xdp_rxq_info *rxq;
- struct xdp_txq_info *txq;
+ union {
+ /* Valid for DEVMAP egress programs. */
+ struct xdp_txq_info *txq;
+ /* Valid for receive buffers backed by non-page netmem. */
+ netmem_ref netmem;
+ };
union {
struct {
@@ -104,6 +112,8 @@ struct xdp_buff {
u64 frame_sz_flags_init;
#endif
};
+ /* Kernel-owned fragment metadata for non-page netmem. */
+ struct skb_shared_info *sinfo;
};
static __always_inline void xdp_reinit_buff(struct xdp_buff *xdp)
@@ -138,7 +148,7 @@ static __always_inline void xdp_buff_set_frag_unreadable(struct xdp_buff *xdp)
static __always_inline u32 xdp_buff_get_skb_flags(const struct xdp_buff *xdp)
{
- return xdp->flags;
+ return xdp->flags & ~XDP_FLAGS_HAS_NETMEM;
}
static __always_inline void xdp_buff_clear_frag_pfmemalloc(struct xdp_buff *xdp)
@@ -150,6 +160,10 @@ static __always_inline void
xdp_init_buff(struct xdp_buff *xdp, u32 frame_sz, struct xdp_rxq_info *rxq)
{
xdp->rxq = rxq;
+ /* Do not initialize the txq/netmem union here. DEVMAP generic XDP
+ * sets txq before this helper is called; receive drivers set netmem
+ * explicitly.
+ */
#ifdef __LITTLE_ENDIAN
/*
@@ -176,6 +190,33 @@ xdp_prepare_buff(struct xdp_buff *xdp, unsigned char *hard_start,
xdp->data_meta = meta_valid ? data : data + 1;
}
+static __always_inline bool xdp_buff_has_netmem(const struct xdp_buff *xdp)
+{
+ return !!(xdp->flags & XDP_FLAGS_HAS_NETMEM);
+}
+
+static __always_inline netmem_ref
+xdp_buff_get_netmem(const struct xdp_buff *xdp)
+{
+ if (xdp_buff_has_netmem(xdp))
+ return xdp->netmem;
+
+ return virt_to_netmem(xdp->data);
+}
+
+static __always_inline void
+xdp_init_buff_from_netmem(struct xdp_buff *xdp, u32 frame_sz,
+ struct xdp_rxq_info *rxq, netmem_ref netmem,
+ struct skb_shared_info *sinfo)
+{
+ xdp_init_buff(xdp, frame_sz, rxq);
+ if (netmem_is_net_iov(netmem)) {
+ xdp->netmem = netmem;
+ xdp->flags |= XDP_FLAGS_HAS_NETMEM;
+ xdp->sinfo = sinfo;
+ }
+}
+
/* Reserve memory area at end-of data area.
*
* This macro reserves tailroom in the XDP buffer by limiting the
@@ -189,6 +230,9 @@ xdp_prepare_buff(struct xdp_buff *xdp, unsigned char *hard_start,
static inline struct skb_shared_info *
xdp_get_shared_info_from_buff(const struct xdp_buff *xdp)
{
+ if (xdp_buff_has_netmem(xdp))
+ return xdp->sinfo;
+
return (struct skb_shared_info *)xdp_data_hard_end(xdp);
}
@@ -290,7 +334,8 @@ static inline bool xdp_buff_add_frag(struct xdp_buff *xdp, netmem_ref netmem,
if (unlikely(netmem_is_pfmemalloc(netmem)))
xdp_buff_set_frag_pfmemalloc(xdp);
- if (unlikely(netmem_is_net_iov(netmem)))
+ if (unlikely(netmem_is_net_iov(netmem) &&
+ !net_iov_is_readable(netmem_to_net_iov(netmem))))
xdp_buff_set_frag_unreadable(xdp);
return true;
@@ -425,7 +470,7 @@ int xdp_update_frame_from_buff(const struct xdp_buff *xdp,
xdp_frame->headroom = headroom - sizeof(*xdp_frame);
xdp_frame->metasize = metasize;
xdp_frame->frame_sz = xdp->frame_sz;
- xdp_frame->flags = xdp->flags;
+ xdp_frame->flags = xdp_buff_get_skb_flags(xdp);
return 0;
}
@@ -436,7 +481,8 @@ struct xdp_frame *xdp_convert_buff_to_frame(struct xdp_buff *xdp)
{
struct xdp_frame *xdp_frame;
- if (xdp->rxq->mem.type == MEM_TYPE_XSK_BUFF_POOL)
+ if (xdp->rxq->mem.type == MEM_TYPE_XSK_BUFF_POOL ||
+ xdp_buff_has_netmem(xdp))
return xdp_convert_zc_to_xdp_frame(xdp);
/* Store info in top of packet */
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 07/15] xsk: Keep the DMA mapping in the buffer pool
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (5 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 06/15] xdp: Track non-page netmem in receive buffers Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 08/15] xsk: Handle a detached FILL ring in RX wakeup Björn Töpel
` (7 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
At unmap, an XSK buffer pool looks up its DMA mapping again. It
searches the UMEM's mapping list for the pool's netdev. A memory
provider, added later in this series, can keep the mapping after the
queue is gone. By then the netdev may be cleared, so the search is
not reliable.
Store the mapping that xp_dma_map() picked in the pool, and unmap
that mapping. Hold a reference on the DMA device while the mapping
exists, so a late final unmap does not use a freed device.
The reference counting of shared mappings does not change.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/xsk_buff_pool.h | 1 +
net/xdp/xsk_buff_pool.c | 8 +++++---
2 files changed, 6 insertions(+), 3 deletions(-)
diff --git a/include/net/xsk_buff_pool.h b/include/net/xsk_buff_pool.h
index a7df573784fd..77264c4902c0 100644
--- a/include/net/xsk_buff_pool.h
+++ b/include/net/xsk_buff_pool.h
@@ -67,6 +67,7 @@ struct xsk_buff_pool {
* even when they are identical.
*/
dma_addr_t *dma_pages;
+ struct xsk_dma_map *dma_map;
struct xdp_buff_xsk *heads;
struct xdp_desc *tx_descs;
u64 chunk_mask;
diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
index f01e1de7360e..244776a72961 100644
--- a/net/xdp/xsk_buff_pool.c
+++ b/net/xdp/xsk_buff_pool.c
@@ -373,7 +373,7 @@ static struct xsk_dma_map *xp_create_dma_map(struct device *dev, struct net_devi
}
dma_map->netdev = netdev;
- dma_map->dev = dev;
+ dma_map->dev = get_device(dev);
dma_map->dma_pages_cnt = nr_pages;
refcount_set(&dma_map->users, 1);
list_add(&dma_map->list, &umem->xsk_dma_list);
@@ -384,6 +384,7 @@ static void xp_destroy_dma_map(struct xsk_dma_map *dma_map)
{
list_del(&dma_map->list);
kvfree(dma_map->dma_pages);
+ put_device(dma_map->dev);
kfree(dma_map);
}
@@ -407,12 +408,11 @@ static void __xp_dma_unmap(struct xsk_dma_map *dma_map, unsigned long attrs)
void xp_dma_unmap(struct xsk_buff_pool *pool, unsigned long attrs)
{
- struct xsk_dma_map *dma_map;
+ struct xsk_dma_map *dma_map = pool->dma_map;
if (!pool->dma_pages)
return;
- dma_map = xp_find_dma_map(pool);
if (!dma_map) {
WARN(1, "Could not find dma_map for device");
return;
@@ -424,6 +424,7 @@ void xp_dma_unmap(struct xsk_buff_pool *pool, unsigned long attrs)
kvfree(pool->dma_pages);
pool->dma_pages = NULL;
pool->dma_pages_cnt = 0;
+ pool->dma_map = NULL;
pool->dev = NULL;
}
EXPORT_SYMBOL(xp_dma_unmap);
@@ -460,6 +461,7 @@ static int xp_init_dma_info(struct xsk_buff_pool *pool, struct xsk_dma_map *dma_
return -ENOMEM;
pool->dev = dma_map->dev;
+ pool->dma_map = dma_map;
pool->dma_pages_cnt = dma_map->dma_pages_cnt;
memcpy(pool->dma_pages, dma_map->dma_pages,
pool->dma_pages_cnt * sizeof(*pool->dma_pages));
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 08/15] xsk: Handle a detached FILL ring in RX wakeup
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (6 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 07/15] xsk: Keep the DMA mapping in the buffer pool Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 09/15] xsk: Add a page-pool memory provider for UMEM Björn Töpel
` (6 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
A buffer pool used as a memory provider can live longer than its
FILL ring, while an old page pool finishes a deferred destroy. The
RX need-wakeup helpers use pool->fq without a NULL check.
Read pool->fq once and return if it is NULL. A normal pool always
has a FILL ring here, so nothing changes for it.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
net/xdp/xsk.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/net/xdp/xsk.c b/net/xdp/xsk.c
index 33475b180ea6..b68dda9c37d1 100644
--- a/net/xdp/xsk.c
+++ b/net/xdp/xsk.c
@@ -50,10 +50,15 @@ static struct kmem_cache *xsk_tx_generic_cache;
void xsk_set_rx_need_wakeup(struct xsk_buff_pool *pool)
{
+ struct xsk_queue *fq = READ_ONCE(pool->fq);
+
+ if (!fq)
+ return;
+
if (pool->cached_need_wakeup & XDP_WAKEUP_RX)
return;
- pool->fq->ring->flags |= XDP_RING_NEED_WAKEUP;
+ fq->ring->flags |= XDP_RING_NEED_WAKEUP;
pool->cached_need_wakeup |= XDP_WAKEUP_RX;
}
EXPORT_SYMBOL(xsk_set_rx_need_wakeup);
@@ -77,10 +82,15 @@ EXPORT_SYMBOL(xsk_set_tx_need_wakeup);
void xsk_clear_rx_need_wakeup(struct xsk_buff_pool *pool)
{
+ struct xsk_queue *fq = READ_ONCE(pool->fq);
+
+ if (!fq)
+ return;
+
if (!(pool->cached_need_wakeup & XDP_WAKEUP_RX))
return;
- pool->fq->ring->flags &= ~XDP_RING_NEED_WAKEUP;
+ fq->ring->flags &= ~XDP_RING_NEED_WAKEUP;
pool->cached_need_wakeup &= ~XDP_WAKEUP_RX;
}
EXPORT_SYMBOL(xsk_clear_rx_need_wakeup);
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 09/15] xsk: Add a page-pool memory provider for UMEM
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (7 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 08/15] xsk: Handle a detached FILL ring in RX wakeup Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 10/15] xsk: Add RX helpers for page-pool drivers Björn Töpel
` (5 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
AF_XDP zero-copy drivers get RX buffers from an xsk_buff_pool. A
driver built on page_pool would need a second RX allocator for that.
Instead, let an XSK buffer pool act as a page-pool memory provider.
It hands out UMEM chunks as NET_IOV_XSK net_iovs, and the driver
uses the normal page_pool API.
A driver calls xsk_pool_setup_page_pool() for XDP_SETUP_XSK_POOL.
For now only 4 KiB pages and aligned UMEM with 4 KiB chunks work;
other setups get -EOPNOTSUPP. A bind without XDP_ZEROCOPY then falls
back to copy mode, as before. The exception is a failed setup that
leaves an old page pool being destroyed. That page pool still uses
the DMA mapping, so the bind fails.
The provider reads the FILL ring in batches of the page-pool cache
refill size, and it handles RX need-wakeup. Each allocation reads at
most one batch, which limits the work spent on bad descriptors in
NAPI. Addresses outside the UMEM, and addresses that are already in
use, count as invalid descriptors and are dropped. refill_done keeps
NAPI scheduled while the FILL ring has entries. When the ring is
empty, it sets NEED_WAKEUP and then checks the ring once more.
The provider asks for page-sized buffers with the UMEM headroom plus
XDP_PACKET_HEADROOM. It refuses a page pool whose DMA sync range
goes past the chunk.
During a queue restart, two page pools can use one provider at the
same time. A provider lock protects the FILL ring, the reuse stack
and need-wakeup. Allocation takes it once per page-pool cache refill
of up to 64 buffers. A packet that is copied into a socket with a
provider takes it once per packet, because the copy also reads the
FILL ring.
A buffer has one owner at a time, as a normal page-pool page does.
Allocation claims a buffer by setting its page pool under the
provider lock. Release clears the link, so destroying the provider
does not need to scan the UMEM.
Generic XDP and synthetic RX queues, such as CPUMAP, would write to
the socket RX ring outside the queue's NAPI. While the provider is
installed, generic XDP drops such packets and counts them in
rx_dropped, and synthetic queues get -EINVAL. Classic zero-copy
sockets do not change.
A failed queue restart can leave the old page pool draining. Keep
the UMEM, the DMA mapping and the netdev until the provider's last
page pool is destroyed. Charge the provider arrays, whose size grows
with the UMEM, to the memory cgroup of the socket owner.
XDP_SOCKETS now selects PAGE_POOL.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/xdp_sock_drv.h | 12 +
include/net/xsk_buff_pool.h | 6 +-
net/core/page_pool.c | 4 +-
net/xdp/Kconfig | 1 +
net/xdp/xsk.c | 50 ++-
net/xdp/xsk.h | 58 ++++
net/xdp/xsk_buff_pool.c | 654 +++++++++++++++++++++++++++++++++++-
7 files changed, 756 insertions(+), 29 deletions(-)
diff --git a/include/net/xdp_sock_drv.h b/include/net/xdp_sock_drv.h
index d94aeb506379..b9288f5dd48b 100644
--- a/include/net/xdp_sock_drv.h
+++ b/include/net/xdp_sock_drv.h
@@ -9,6 +9,8 @@
#include <net/xdp_sock.h>
#include <net/xsk_buff_pool.h>
+struct netlink_ext_ack;
+
#define XDP_UMEM_MIN_CHUNK_SHIFT 11
#define XDP_UMEM_MIN_CHUNK_SIZE (1 << XDP_UMEM_MIN_CHUNK_SHIFT)
@@ -28,6 +30,8 @@ void xsk_tx_completed(struct xsk_buff_pool *pool, u32 nb_entries);
bool xsk_tx_peek_desc(struct xsk_buff_pool *pool, struct xdp_desc *desc);
u32 xsk_tx_peek_release_desc_batch(struct xsk_buff_pool *pool, u32 max);
void xsk_tx_release(struct xsk_buff_pool *pool);
+int xsk_pool_setup_page_pool(struct net_device *dev, struct xsk_buff_pool *pool,
+ u16 queue_id, struct netlink_ext_ack *extack);
struct xsk_buff_pool *xsk_get_pool_from_qid(struct net_device *dev,
u16 queue_id);
void xsk_set_rx_need_wakeup(struct xsk_buff_pool *pool);
@@ -370,6 +374,14 @@ static inline void xsk_tx_release(struct xsk_buff_pool *pool)
{
}
+static inline int xsk_pool_setup_page_pool(struct net_device *dev,
+ struct xsk_buff_pool *pool,
+ u16 queue_id,
+ struct netlink_ext_ack *extack)
+{
+ return -EOPNOTSUPP;
+}
+
static inline struct xsk_buff_pool *
xsk_get_pool_from_qid(struct net_device *dev, u16 queue_id)
{
diff --git a/include/net/xsk_buff_pool.h b/include/net/xsk_buff_pool.h
index 77264c4902c0..9eed8796a356 100644
--- a/include/net/xsk_buff_pool.h
+++ b/include/net/xsk_buff_pool.h
@@ -11,6 +11,7 @@
#include <net/xdp.h>
struct xsk_buff_pool;
+struct xsk_pp;
struct xdp_rxq_info;
struct xsk_cb_desc;
struct xsk_queue;
@@ -52,7 +53,8 @@ struct xsk_buff_pool {
spinlock_t xsk_tx_list_lock;
refcount_t users;
struct xdp_umem *umem;
- struct work_struct work;
+ struct xsk_pp *pp;
+ struct delayed_work work;
/* Protects generic receive in shared and non-shared umem mode. */
spinlock_t rx_lock;
struct list_head free_list;
@@ -117,10 +119,8 @@ int xp_alloc_tx_descs(struct xsk_buff_pool *pool, struct xdp_sock *xs,
void xp_destroy(struct xsk_buff_pool *pool);
void xp_get_pool(struct xsk_buff_pool *pool);
bool xp_put_pool(struct xsk_buff_pool *pool);
-void xp_clear_dev(struct xsk_buff_pool *pool);
void xp_add_xsk(struct xsk_buff_pool *pool, struct xdp_sock *xs);
void xp_del_xsk(struct xsk_buff_pool *pool, struct xdp_sock *xs);
-
/* AF_XDP, and XDP core. */
void xp_free(struct xdp_buff_xsk *xskb);
diff --git a/net/core/page_pool.c b/net/core/page_pool.c
index e36ae6123adf..1f446e9a10d7 100644
--- a/net/core/page_pool.c
+++ b/net/core/page_pool.c
@@ -1397,8 +1397,8 @@ void net_mp_release_page_pool_bulk(struct page_pool *pool,
atomic_add(count, release_cnt);
}
-/* Disassociate a niov from a page pool. Should only be used in the
- * ->release_netmem() path.
+/* Disassociate a niov from a page pool. Memory providers may do this either
+ * from ->release_netmem() or from ->destroy() after all objects were released.
*/
void net_mp_niov_clear_page_pool(struct net_iov *niov)
{
diff --git a/net/xdp/Kconfig b/net/xdp/Kconfig
index 71af2febe72a..c9c68d3b3712 100644
--- a/net/xdp/Kconfig
+++ b/net/xdp/Kconfig
@@ -2,6 +2,7 @@
config XDP_SOCKETS
bool "XDP sockets"
depends on BPF_SYSCALL
+ select PAGE_POOL
default n
help
XDP sockets allows a channel between XDP programs and
diff --git a/net/xdp/xsk.c b/net/xdp/xsk.c
index b68dda9c37d1..90c98b42a18a 100644
--- a/net/xdp/xsk.c
+++ b/net/xdp/xsk.c
@@ -464,7 +464,9 @@ static int xsk_rcv_check(struct xdp_sock *xs, struct xdp_buff *xdp, u32 len)
static void xsk_flush(struct xdp_sock *xs)
{
xskq_prod_submit(xs->rx);
- __xskq_cons_release(xs->pool->fq);
+ /* Provider pools publish FILL consumption under the provider lock. */
+ if (!READ_ONCE(xs->pool->pp))
+ __xskq_cons_release(xs->pool->fq);
sock_def_readable(&xs->sk);
}
@@ -474,16 +476,40 @@ int xsk_generic_rcv(struct xdp_sock *xs, struct xdp_buff *xdp)
int err;
err = xsk_rcv_check(xs, xdp, len);
- if (!err) {
- spin_lock_bh(&xs->pool->rx_lock);
- err = __xsk_rcv(xs, xdp, len);
- xsk_flush(xs);
+ if (err)
+ return err;
+ spin_lock_bh(&xs->pool->rx_lock);
+ if (unlikely(READ_ONCE(xs->pool->pp))) {
+ xs->rx_dropped++;
spin_unlock_bh(&xs->pool->rx_lock);
+ return -EOPNOTSUPP;
}
+ err = __xsk_rcv(xs, xdp, len);
+ xsk_flush(xs);
+ spin_unlock_bh(&xs->pool->rx_lock);
return err;
}
+/* Copy into the socket's UMEM. A provider-backed pool shares its FILL ring
+ * with provider allocation, which can run for another page pool of the queue
+ * while the queue is replaced.
+ */
+static int xsk_rcv_copy(struct xdp_sock *xs, struct xdp_buff *xdp, u32 len)
+{
+ struct xsk_pp *provider = READ_ONCE(xs->pool->pp);
+ int err;
+
+ if (likely(!provider))
+ return __xsk_rcv(xs, xdp, len);
+
+ spin_lock_bh(&provider->lock);
+ err = __xsk_rcv(xs, xdp, len);
+ __xskq_cons_release(xs->pool->fq);
+ spin_unlock_bh(&provider->lock);
+ return err;
+}
+
static int xsk_rcv(struct xdp_sock *xs, struct xdp_buff *xdp)
{
u32 len = xdp_get_buff_len(xdp);
@@ -498,7 +524,15 @@ static int xsk_rcv(struct xdp_sock *xs, struct xdp_buff *xdp)
return xsk_rcv_zc(xs, xdp, len);
}
- err = __xsk_rcv(xs, xdp, len);
+ /* The socket RX ring has a single producer, the queue's poll context.
+ * Reject synthetic RX queues before their remote context produces
+ * into a provider-backed socket.
+ */
+ if (unlikely(READ_ONCE(xs->pool->pp) &&
+ !xdp_rxq_info_is_reg(xdp->rxq)))
+ return -EINVAL;
+
+ err = xsk_rcv_copy(xs, xdp, len);
if (!err)
xdp_return_buff(xdp);
return err;
@@ -2149,8 +2183,8 @@ static int xsk_notifier(struct notifier_block *this,
xsk_unbind_dev(xs);
- /* Clear device references. */
- xp_clear_dev(xs->pool);
+ /* Unregister cannot hold a device reference. */
+ xp_clear_dev(xs->pool, XSK_POOL_CLEAR_FORCE);
}
mutex_unlock(&xs->mutex);
}
diff --git a/net/xdp/xsk.h b/net/xdp/xsk.h
index 7c811b5cce76..8770778cd322 100644
--- a/net/xdp/xsk.h
+++ b/net/xdp/xsk.h
@@ -4,6 +4,64 @@
#ifndef XSK_H_
#define XSK_H_
+#include <net/netmem.h>
+
+struct xsk_buff_pool;
+
+enum xsk_pool_clear_mode {
+ XSK_POOL_CLEAR_NORMAL,
+ XSK_POOL_CLEAR_FORCE,
+};
+
+struct xsk_pp_info {
+ struct net_iov_area area;
+ struct xsk_buff_pool *pool;
+ u32 chunk_shift;
+};
+
+/* Keep page_pool descriptors independent from direct-driver XSK buffers.
+ * Queue replacement creates the new page_pool before it stops the old queue
+ * and destroys the old page_pool after the new queue started, so two
+ * page_pools can use one provider. @lock serializes the state they share.
+ * Generic XDP cannot deliver to a provider-backed socket.
+ */
+struct xsk_pp {
+ struct xsk_pp_info info;
+ struct xsk_queue __rcu *fq;
+ /* FILL consumers, @reuse and RX need_wakeup state */
+ spinlock_t lock;
+ u32 reuse_cnt;
+ u32 nr_pools;
+ u64 chunk_mask;
+ u64 addrs_cnt;
+ u8 release_retries;
+ bool dma_need_sync;
+ bool detached;
+ u32 reuse[];
+};
+
+void xp_clear_dev(struct xsk_buff_pool *pool, enum xsk_pool_clear_mode mode);
+
+static inline bool xp_netmem_is_xsk(netmem_ref netmem)
+{
+ return netmem_is_net_iov(netmem) &&
+ netmem_to_net_iov(netmem)->type == NET_IOV_XSK;
+}
+
+static inline struct xsk_pp_info *xp_netmem_to_pp(netmem_ref netmem)
+{
+ struct net_iov *niov = netmem_to_net_iov(netmem);
+
+ return container_of(net_iov_owner(niov), struct xsk_pp_info, area);
+}
+
+static inline bool xp_netmem_is_from_pool(netmem_ref netmem,
+ const struct xsk_buff_pool *pool)
+{
+ return xp_netmem_is_xsk(netmem) &&
+ xp_netmem_to_pp(netmem)->pool == pool;
+}
+
struct xdp_ring_offset_v1 {
__u64 producer;
__u64 consumer;
diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
index 244776a72961..e25347f8c208 100644
--- a/net/xdp/xsk_buff_pool.c
+++ b/net/xdp/xsk_buff_pool.c
@@ -1,7 +1,11 @@
// SPDX-License-Identifier: GPL-2.0
#include <linux/netdevice.h>
+#include <linux/sizes.h>
#include <net/netdev_lock.h>
+#include <net/netdev_queues.h>
+#include <net/netdev_rx_queue.h>
+#include <net/page_pool/helpers.h>
#include <net/page_pool/memory_provider.h>
#include <net/xsk_buff_pool.h>
#include <net/xdp_sock.h>
@@ -12,6 +16,20 @@
#include "xsk.h"
#define ETH_PAD_LEN (ETH_HLEN + 2 * VLAN_HLEN + ETH_FCS_LEN)
+#define XSK_PAGE_POOL_DMA_ATTR (DMA_ATTR_WEAK_ORDERING | \
+ DMA_ATTR_SKIP_CPU_SYNC)
+#define XSK_POOL_RELEASE_RETRY_MAX (60 * HZ)
+
+static bool xp_pp_teardown(struct xsk_buff_pool *pool);
+static void xp_pp_retry_release(struct xsk_buff_pool *pool);
+static void xp_destroy_unbound_deferred(struct work_struct *work);
+
+static void __xp_destroy(struct xsk_buff_pool *pool)
+{
+ kvfree(pool->tx_descs);
+ kvfree(pool->heads);
+ kvfree(pool);
+}
void xp_add_xsk(struct xsk_buff_pool *pool, struct xdp_sock *xs)
{
@@ -38,9 +56,22 @@ void xp_destroy(struct xsk_buff_pool *pool)
if (!pool)
return;
- kvfree(pool->tx_descs);
- kvfree(pool->heads);
- kvfree(pool);
+ /* A failed queue replacement can leave a page_pool waiting for an
+ * in-flight buffer. Keep the UMEM and pool alive until its provider
+ * destroy callback has run.
+ */
+ if (pool->pp) {
+ rcu_assign_pointer(pool->pp->fq, NULL);
+ WRITE_ONCE(pool->fq, NULL);
+ WRITE_ONCE(pool->cq, NULL);
+ synchronize_net();
+ xdp_get_umem(pool->umem);
+ INIT_DELAYED_WORK(&pool->work, xp_destroy_unbound_deferred);
+ schedule_delayed_work(&pool->work, 0);
+ return;
+ }
+
+ __xp_destroy(pool);
}
int xp_alloc_tx_descs(struct xsk_buff_pool *pool, struct xdp_sock *xs,
@@ -167,6 +198,7 @@ int xp_assign_dev(struct xsk_buff_pool *pool,
{
u32 needed = netdev->mtu + ETH_PAD_LEN;
u32 segs = netdev->xdp_zc_max_segs;
+ bool old_zc = pool->umem->zc;
bool mbuf = flags & XDP_USE_SG;
bool force_zc, force_copy;
struct netdev_bpf bpf;
@@ -250,23 +282,35 @@ int xp_assign_dev(struct xsk_buff_pool *pool,
if (err)
goto err_unreg_pool;
+ /* Record the successful driver install before validating its result so
+ * the error path can issue the matching XDP_SETUP_XSK_POOL teardown.
+ */
+ pool->umem->zc = true;
if (!pool->dma_pages) {
WARN(1, "Driver did not DMA map zero-copy buffers");
err = -EINVAL;
goto err_unreg_xsk;
}
- pool->umem->zc = true;
pool->xdp_zc_max_segs = netdev->xdp_zc_max_segs;
return 0;
err_unreg_xsk:
xp_disable_drv_zc(pool);
+ if (pool->pp)
+ pool->pp->detached = true;
+ pool->umem->zc = old_zc;
err_unreg_pool:
- if (!force_zc)
+ /* A provider whose page_pool destruction was deferred still owns the
+ * DMA mapping. Do not turn that failed zero-copy setup into copy mode.
+ */
+ if (!force_zc && !pool->pp)
err = 0; /* fallback to copy mode */
if (err) {
xsk_clear_pool_at_qid(netdev, queue_id);
- dev_put(netdev);
+ if (!pool->pp) {
+ pool->netdev = NULL;
+ dev_put(netdev);
+ }
}
return err;
}
@@ -288,29 +332,75 @@ int xp_assign_dev_shared(struct xsk_buff_pool *pool, struct xdp_sock *umem_xs,
return xp_assign_dev(pool, dev, queue_id, flags);
}
-void xp_clear_dev(struct xsk_buff_pool *pool)
+static void xp_pp_force_dma_unmap(struct xsk_buff_pool *pool)
+{
+ if (!pool->dma_map)
+ return;
+
+ pool->dma_map->netdev = NULL;
+ xsk_pool_dma_unmap(pool, XSK_PAGE_POOL_DMA_ATTR);
+}
+
+void xp_clear_dev(struct xsk_buff_pool *pool, enum xsk_pool_clear_mode mode)
{
struct net_device *netdev = pool->netdev;
+ struct xsk_pp *provider = pool->pp;
- if (!pool->netdev)
+ if (!netdev)
return;
+ /* Keep this reference after a normal detach. It keeps the netdev
+ * identity stable for shared DMA-map lookup and lets unregister
+ * invalidate that mapping while the old page_pool drains.
+ */
+ if (provider && provider->detached) {
+ if (mode != XSK_POOL_CLEAR_FORCE)
+ return;
+ xp_pp_force_dma_unmap(pool);
+ pool->netdev = NULL;
+ dev_put(netdev);
+ return;
+ }
netdev_lock_ops(netdev);
xp_disable_drv_zc(pool);
xsk_clear_pool_at_qid(pool->netdev, pool->queue_id);
- pool->netdev = NULL;
+ if (provider) {
+ provider->detached = true;
+ if (mode == XSK_POOL_CLEAR_FORCE) {
+ xp_pp_force_dma_unmap(pool);
+ pool->netdev = NULL;
+ }
+ } else {
+ pool->netdev = NULL;
+ }
netdev_unlock_ops(netdev);
- dev_put(netdev);
+
+ if (!provider || mode == XSK_POOL_CLEAR_FORCE)
+ dev_put(netdev);
}
static void xp_release_deferred(struct work_struct *work)
{
- struct xsk_buff_pool *pool = container_of(work, struct xsk_buff_pool,
- work);
+ struct net_device *netdev = NULL;
+ struct xsk_buff_pool *pool;
+ bool teardown_done;
+
+ pool = container_of(to_delayed_work(work), struct xsk_buff_pool, work);
rtnl_lock();
- xp_clear_dev(pool);
+ xp_clear_dev(pool, XSK_POOL_CLEAR_NORMAL);
+ teardown_done = xp_pp_teardown(pool);
+ if (teardown_done && pool->netdev) {
+ netdev = pool->netdev;
+ pool->netdev = NULL;
+ }
rtnl_unlock();
+ if (!teardown_done) {
+ xp_pp_retry_release(pool);
+ return;
+ }
+ if (netdev)
+ dev_put(netdev);
if (pool->fq) {
xskq_destroy(pool->fq);
@@ -323,7 +413,7 @@ static void xp_release_deferred(struct work_struct *work)
}
xdp_put_umem(pool->umem, false);
- xp_destroy(pool);
+ __xp_destroy(pool);
}
void xp_get_pool(struct xsk_buff_pool *pool)
@@ -337,8 +427,8 @@ bool xp_put_pool(struct xsk_buff_pool *pool)
return false;
if (refcount_dec_and_test(&pool->users)) {
- INIT_WORK(&pool->work, xp_release_deferred);
- schedule_work(&pool->work);
+ INIT_DELAYED_WORK(&pool->work, xp_release_deferred);
+ schedule_delayed_work(&pool->work, 0);
return true;
}
@@ -514,6 +604,538 @@ int xp_dma_map(struct xsk_buff_pool *pool, struct device *dev,
}
EXPORT_SYMBOL(xp_dma_map);
+static void xp_pp_free(struct xsk_buff_pool *pool)
+{
+ struct xsk_pp *provider = pool->pp;
+ u32 idx;
+
+ if (!provider)
+ return;
+ if (WARN_ON_ONCE(READ_ONCE(provider->nr_pools)))
+ return;
+
+ /* A failed queue install may have consumed FILL entries before the
+ * replacement queue failed to start. Preserve those frames for the
+ * direct or copy-mode allocator selected by bind fallback.
+ */
+ while (provider->reuse_cnt) {
+ idx = provider->reuse[--provider->reuse_cnt];
+ xp_free(&pool->heads[idx]);
+ }
+
+ spin_lock_bh(&pool->rx_lock);
+ WRITE_ONCE(pool->pp, NULL);
+ spin_unlock_bh(&pool->rx_lock);
+ kvfree(provider->info.area.niovs);
+ kvfree(provider);
+}
+
+static bool xp_pp_teardown(struct xsk_buff_pool *pool)
+{
+ struct xsk_pp *provider = pool->pp;
+ u32 busy;
+
+ if (!provider)
+ return true;
+ if (WARN_ON_ONCE(!provider->detached))
+ return false;
+ /* xp_pp_destroy() touches the provider until it drops the lock. */
+ spin_lock_bh(&provider->lock);
+ busy = provider->nr_pools;
+ spin_unlock_bh(&provider->lock);
+ if (busy)
+ return false;
+
+ xsk_pool_dma_unmap(pool, XSK_PAGE_POOL_DMA_ATTR);
+ xp_pp_free(pool);
+
+ return true;
+}
+
+static void xp_pp_retry_release(struct xsk_buff_pool *pool)
+{
+ struct xsk_pp *provider = pool->pp;
+ unsigned long delay = HZ;
+ u8 retries;
+
+ if (provider) {
+ retries = provider->release_retries;
+ if (retries < 6)
+ provider->release_retries++;
+ delay = min_t(unsigned long, HZ << retries,
+ XSK_POOL_RELEASE_RETRY_MAX);
+ }
+ schedule_delayed_work(&pool->work, delay);
+}
+
+static void xp_destroy_unbound_deferred(struct work_struct *work)
+{
+ struct xsk_buff_pool *pool;
+ struct net_device *netdev;
+
+ pool = container_of(to_delayed_work(work), struct xsk_buff_pool, work);
+
+ rtnl_lock();
+ if (!xp_pp_teardown(pool)) {
+ rtnl_unlock();
+ xp_pp_retry_release(pool);
+ return;
+ }
+ netdev = pool->netdev;
+ pool->netdev = NULL;
+ rtnl_unlock();
+
+ if (netdev)
+ dev_put(netdev);
+ xdp_put_umem(pool->umem, false);
+ __xp_destroy(pool);
+}
+
+static struct xsk_pp *xp_pp_create(struct xsk_buff_pool *pool)
+{
+ struct xsk_pp *provider;
+ struct net_iov *niov;
+ dma_addr_t dma;
+ u64 addr;
+ u32 i;
+
+ /* The provider arrays scale with the UMEM; charge them like it. */
+ provider = kvzalloc_flex(*provider, reuse, pool->heads_cnt,
+ GFP_KERNEL_ACCOUNT);
+ if (!provider)
+ return ERR_PTR(-ENOMEM);
+
+ provider->info.area.niovs =
+ kvzalloc_objs(*provider->info.area.niovs, pool->heads_cnt,
+ GFP_KERNEL_ACCOUNT);
+ if (!provider->info.area.niovs) {
+ kvfree(provider);
+ return ERR_PTR(-ENOMEM);
+ }
+ provider->info.pool = pool;
+ RCU_INIT_POINTER(provider->fq, pool->fq);
+ spin_lock_init(&provider->lock);
+ provider->info.area.num_niovs = pool->heads_cnt;
+ provider->info.area.vaddr = pool->addrs;
+ provider->info.area.niov_shift = pool->chunk_shift;
+ provider->chunk_mask = pool->chunk_mask;
+ provider->addrs_cnt = pool->addrs_cnt;
+ provider->info.chunk_shift = pool->chunk_shift;
+ provider->dma_need_sync = dma_dev_need_sync(pool->dev);
+
+ for (i = 0; i < provider->info.area.num_niovs; i++) {
+ niov = &provider->info.area.niovs[i];
+ addr = (u64)i << provider->info.chunk_shift;
+ dma = (pool->dma_pages[addr >> PAGE_SHIFT] &
+ ~XSK_NEXT_PG_CONTIG_MASK) + (addr & ~PAGE_MASK);
+
+ net_iov_init(niov, &provider->info.area, NET_IOV_XSK);
+ if (net_mp_niov_set_dma_addr(niov, dma))
+ goto err_free_niovs;
+ }
+
+ /* Exclude generic receive. It can run on another CPU after RPS and
+ * would produce into the socket RX ring, which provider delivery
+ * fills without a lock from the queue's poll context.
+ */
+ spin_lock_bh(&pool->rx_lock);
+ WRITE_ONCE(pool->pp, provider);
+ spin_unlock_bh(&pool->rx_lock);
+ return provider;
+
+err_free_niovs:
+ kvfree(provider->info.area.niovs);
+ kvfree(provider);
+ return ERR_PTR(-ERANGE);
+}
+
+static int xp_pp_init(struct page_pool *pp)
+{
+ struct xsk_buff_pool *pool = pp->mp_priv;
+ struct xsk_pp *provider = pool->pp;
+ int err = 0;
+
+ if (!provider || !pool->dma_pages)
+ return -EINVAL;
+ if (pp->p.order)
+ return -E2BIG;
+ if (pp->p.dev != pool->dev ||
+ pp->p.dma_dir != DMA_BIDIRECTIONAL)
+ return -EINVAL;
+ /* Each object is one UMEM chunk. Syncs and device writes described by
+ * the page_pool must stay inside it.
+ */
+ if (pp->p.offset + pp->p.max_len > pool->chunk_size)
+ return -EINVAL;
+
+ spin_lock_bh(&provider->lock);
+ if (provider->detached)
+ err = -ENODEV;
+ else
+ provider->nr_pools++;
+ spin_unlock_bh(&provider->lock);
+
+ return err;
+}
+
+static void xp_pp_destroy(struct page_pool *pp)
+{
+ struct xsk_buff_pool *pool = pp->mp_priv;
+ struct xsk_pp *provider = pool->pp;
+
+ /* Every object the page_pool held was released through
+ * xp_pp_release_netmem() or handed to userspace, both of which clear
+ * its page_pool association.
+ */
+ spin_lock_bh(&provider->lock);
+ if (!WARN_ON_ONCE(!provider->nr_pools))
+ provider->nr_pools--;
+ spin_unlock_bh(&provider->lock);
+}
+
+static netmem_ref xp_pp_prepare_netmem(struct page_pool *pp,
+ struct xsk_pp *provider,
+ struct net_iov *niov)
+{
+ netmem_ref netmem = net_iov_to_netmem(niov);
+ dma_addr_t dma = page_pool_get_dma_addr_netmem(netmem);
+
+ if (provider->dma_need_sync)
+ dma_sync_single_range_for_device(pp->p.dev, dma,
+ pp->p.offset, pp->p.max_len,
+ pp->p.dma_dir);
+
+ return netmem;
+}
+
+/* A set pp marks an object a page pool owns. Repeated FILL addresses must
+ * not hand it out twice.
+ */
+static bool xp_pp_claim(struct page_pool *pp, struct net_iov *niov)
+{
+ if (unlikely(niov->desc.pp))
+ return false;
+
+ niov->desc.pp = pp;
+ return true;
+}
+
+static unsigned int
+xp_pp_alloc_reused(struct page_pool *pp, struct xsk_pp *provider,
+ netmem_ref *netmems, unsigned int max)
+{
+ unsigned int entries;
+ unsigned int allocated = 0;
+ struct net_iov *niov;
+ unsigned int i;
+ u32 idx;
+
+ entries = min(max, provider->reuse_cnt);
+ for (i = 0; i < entries; i++) {
+ idx = provider->reuse[--provider->reuse_cnt];
+ niov = &provider->info.area.niovs[idx];
+ if (xp_pp_claim(pp, niov))
+ netmems[allocated++] = net_iov_to_netmem(niov);
+ }
+
+ /* page_pool normally synced these buffers before ->release_netmem(),
+ * but page_pool teardown disables DMA sync before emptying its caches.
+ * Conservatively prepare both reuse and FQ buffers below.
+ */
+ return allocated;
+}
+
+static unsigned int xp_pp_alloc_fq(struct page_pool *pp,
+ struct xsk_pp *provider,
+ netmem_ref *netmems, unsigned int max)
+{
+ struct xsk_buff_pool *pool = provider->info.pool;
+ struct xsk_queue *fq;
+ struct net_iov *niov;
+ u64 addr;
+ unsigned int allocated = 0;
+ u32 cached_cons, entries;
+ u32 cons, idx;
+ u32 i;
+
+ rcu_read_lock_bh();
+ fq = rcu_dereference_bh(provider->fq);
+ if (!fq)
+ goto out;
+
+ entries = xskq_cons_nb_entries(fq, max);
+ if (!entries && pool->uses_need_wakeup) {
+ /* Publish NEED_WAKEUP before checking again. */
+ xsk_set_rx_need_wakeup(pool);
+ /* Pair the publication with the producer recheck. */
+ smp_mb();
+ entries = xskq_cons_nb_entries(fq, max);
+ }
+ if (!entries) {
+ fq->queue_empty_descs++;
+ goto out;
+ }
+
+ cached_cons = fq->cached_cons;
+ for (i = 0; i < entries; i++) {
+ cons = cached_cons++;
+ __xskq_cons_read_addr_unchecked(fq, cons, &addr);
+ addr &= provider->chunk_mask;
+ if (unlikely(addr >= provider->addrs_cnt)) {
+ fq->invalid_descs++;
+ continue;
+ }
+
+ idx = addr >> provider->info.chunk_shift;
+ niov = &provider->info.area.niovs[idx];
+ if (unlikely(!xp_pp_claim(pp, niov))) {
+ fq->invalid_descs++;
+ continue;
+ }
+ netmems[allocated++] = net_iov_to_netmem(niov);
+ }
+
+ xskq_cons_release_n(fq, entries);
+ __xskq_cons_release(fq);
+out:
+ rcu_read_unlock_bh();
+ return allocated;
+}
+
+static netmem_ref xp_pp_alloc_netmems(struct page_pool *pp, gfp_t gfp)
+{
+ struct xsk_buff_pool *pool = pp->mp_priv;
+ struct xsk_pp *provider = pool->pp;
+ netmem_ref *netmems = pp->alloc.cache;
+ unsigned int allocated;
+ unsigned int i;
+
+ if (WARN_ON_ONCE(pp->alloc.count))
+ return 0;
+
+ /* One lock round trip refills a whole page_pool allocation cache. */
+ spin_lock_bh(&provider->lock);
+ allocated = xp_pp_alloc_reused(pp, provider, netmems,
+ PP_ALLOC_CACHE_REFILL);
+ if (allocated < PP_ALLOC_CACHE_REFILL)
+ allocated += xp_pp_alloc_fq(pp, provider, netmems + allocated,
+ PP_ALLOC_CACHE_REFILL - allocated);
+ spin_unlock_bh(&provider->lock);
+ if (!allocated)
+ return 0;
+
+ /* The selected netmems now belong to @pp alone. Finish DMA and
+ * page-pool setup without the lock before the driver sees them.
+ */
+ for (i = 0; i < allocated; i++) {
+ struct net_iov *niov = netmem_to_net_iov(netmems[i]);
+
+ netmems[i] = xp_pp_prepare_netmem(pp, provider, niov);
+ }
+ net_mp_netmem_set_page_pool_bulk(pp, netmems, allocated);
+
+ allocated--;
+ pp->alloc.count = allocated;
+ return netmems[allocated];
+}
+
+static bool xp_pp_release_netmem(struct page_pool *pp, netmem_ref netmem)
+{
+ struct xsk_buff_pool *pool = pp->mp_priv;
+ struct xsk_pp *provider = pool->pp;
+ struct net_iov *niov = netmem_to_net_iov(netmem);
+ u32 idx = net_iov_idx(niov);
+
+ /* page_pool calls this on ptr_ring overflow and from its destroy and
+ * release-retry paths, which can run while another page_pool of this
+ * provider allocates. Clear the association before the object becomes
+ * visible to that page_pool.
+ */
+ net_mp_niov_clear_page_pool(niov);
+ spin_lock_bh(&provider->lock);
+ if (likely(provider->reuse_cnt < provider->info.area.num_niovs))
+ provider->reuse[provider->reuse_cnt++] = idx;
+ spin_unlock_bh(&provider->lock);
+
+ return false;
+}
+
+static bool xp_pp_refill_done(struct page_pool *pp, bool full)
+{
+ struct xsk_buff_pool *pool = pp->mp_priv;
+ struct xsk_pp *provider = pool->pp;
+ bool need_wakeup = pool->uses_need_wakeup;
+ struct xsk_queue *fq;
+ bool pending;
+
+ if (likely(full) &&
+ (!need_wakeup || !(pool->cached_need_wakeup & XDP_WAKEUP_RX)))
+ return true;
+
+ spin_lock_bh(&provider->lock);
+ fq = rcu_dereference_bh(provider->fq);
+ if (!fq) {
+ spin_unlock_bh(&provider->lock);
+ return true;
+ }
+
+ if (full) {
+ if (need_wakeup)
+ xsk_clear_rx_need_wakeup(pool);
+ spin_unlock_bh(&provider->lock);
+ return true;
+ }
+
+ pending = xskq_cons_nb_entries(fq, 1);
+ if (!pending && need_wakeup) {
+ xsk_set_rx_need_wakeup(pool);
+ smp_mb(); /* Order NEED_WAKEUP before the producer recheck. */
+ pending = xskq_cons_nb_entries(fq, 1);
+ }
+ spin_unlock_bh(&provider->lock);
+
+ return pending ? false : need_wakeup;
+}
+
+static int xp_pp_nl_fill(void *mp_priv, struct sk_buff *rsp,
+ struct netdev_rx_queue *rxq)
+{
+ return 0;
+}
+
+static void xp_pp_uninstall(void *mp_priv, struct netdev_rx_queue *rxq)
+{
+ struct xsk_buff_pool *pool = mp_priv;
+ struct xsk_pp *provider = pool->pp;
+ struct net_device *netdev = pool->netdev;
+
+ if (!netdev)
+ goto clear_rxq;
+
+ /* Unregister uninstalls providers before the AF_XDP notifier runs. */
+ xsk_clear_pool_at_qid(netdev, pool->queue_id);
+ if (provider)
+ provider->detached = true;
+ else
+ WARN_ON_ONCE(1);
+ if (pool->dma_map) {
+ pool->dma_map->netdev = NULL;
+ xsk_pool_dma_unmap(pool, XSK_PAGE_POOL_DMA_ATTR);
+ }
+ pool->netdev = NULL;
+ dev_put(netdev);
+
+clear_rxq:
+ /* Existing page_pools retain their private pool pointer until their
+ * deferred destruction; the device queue is no longer operational.
+ */
+ memset(&rxq->mp_params, 0, sizeof(rxq->mp_params));
+}
+
+static const struct memory_provider_ops xsk_pp_ops = {
+ .init = xp_pp_init,
+ .destroy = xp_pp_destroy,
+ .alloc_netmems = xp_pp_alloc_netmems,
+ .release_netmem = xp_pp_release_netmem,
+ .refill_done = xp_pp_refill_done,
+ .nl_fill = xp_pp_nl_fill,
+ .uninstall = xp_pp_uninstall,
+ .caps = MP_CAP_READABLE,
+};
+
+static int xsk_pool_validate_page_pool(struct xsk_buff_pool *pool,
+ struct netlink_ext_ack *extack)
+{
+ if (PAGE_SIZE != SZ_4K) {
+ NL_SET_ERR_MSG(extack,
+ "page pool UMEM requires 4 KiB base pages");
+ return -EOPNOTSUPP;
+ }
+ if (pool->unaligned) {
+ NL_SET_ERR_MSG(extack,
+ "page pool UMEM requires aligned chunks");
+ return -EOPNOTSUPP;
+ }
+ /* Each provider object represents one complete UMEM chunk. */
+ if (pool->chunk_size != PAGE_SIZE) {
+ NL_SET_ERR_MSG(extack,
+ "page pool UMEM requires page-sized chunks");
+ return -EOPNOTSUPP;
+ }
+
+ return 0;
+}
+
+int xsk_pool_setup_page_pool(struct net_device *dev, struct xsk_buff_pool *pool,
+ u16 queue_id, struct netlink_ext_ack *extack)
+{
+ struct xsk_pp *provider;
+ struct device *dma_dev;
+ struct pp_memory_provider_params mp = {
+ .mp_ops = &xsk_pp_ops,
+ };
+ int err;
+
+ ASSERT_RTNL();
+
+ if (!pool) {
+ pool = xsk_get_pool_from_qid(dev, queue_id);
+ if (!pool)
+ return -EINVAL;
+
+ mp.mp_priv = pool;
+ netif_mp_close_rxq(dev, queue_id, &mp);
+ return 0;
+ }
+ if (xsk_get_pool_from_qid(dev, queue_id) != pool) {
+ NL_SET_ERR_MSG(extack,
+ "designated queue has no matching XSK pool");
+ return -EINVAL;
+ }
+ err = xsk_pool_validate_page_pool(pool, extack);
+ if (err)
+ return err;
+
+ dma_dev = netdev_queue_get_dma_dev(dev, queue_id,
+ NETDEV_QUEUE_TYPE_RX);
+ if (!dma_dev) {
+ NL_SET_ERR_MSG(extack, "RX queue has no DMA device");
+ return -EOPNOTSUPP;
+ }
+
+ err = xsk_pool_dma_map(pool, dma_dev, XSK_PAGE_POOL_DMA_ATTR);
+ if (err)
+ return err;
+
+ provider = xp_pp_create(pool);
+ if (IS_ERR(provider)) {
+ err = PTR_ERR(provider);
+ goto err_unmap;
+ }
+
+ mp.mp_priv = pool;
+ mp.rx_page_size = pool->chunk_size;
+ mp.rx_headroom = xsk_pool_get_headroom(pool);
+ err = netif_mp_open_rxq(dev, queue_id, &mp, extack);
+ if (err)
+ goto err_free_provider;
+
+ return 0;
+
+err_free_provider:
+ provider->detached = true;
+ /* A deferred page_pool destroy keeps pool->pp installed. That state
+ * suppresses copy-mode fallback in xp_assign_dev(), while xp_destroy()
+ * keeps the UMEM alive until the provider destroy callback completes.
+ */
+ xp_pp_teardown(pool);
+ return err;
+err_unmap:
+ xsk_pool_dma_unmap(pool, XSK_PAGE_POOL_DMA_ATTR);
+ return err;
+}
+EXPORT_SYMBOL_GPL(xsk_pool_setup_page_pool);
+
static bool xp_addr_crosses_non_contig_pg(struct xsk_buff_pool *pool,
u64 addr)
{
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 10/15] xsk: Add RX helpers for page-pool drivers
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (8 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 09/15] xsk: Add a page-pool memory provider for UMEM Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 11/15] xdp: Copy provider buffers on pass and redirect Björn Töpel
` (4 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
A driver that uses an XSK pool as its page-pool provider must know
whether the pool on a queue is the installed provider, or only
registered on the queue. It also needs the scatter-gather setting of
the pool.
Add xsk_get_pool_from_rxq(), which returns the pool only while it is
the queue's installed provider, and xsk_pool_uses_sg(). Name the
128-byte RX frame alignment XSK_RX_FRAME_SIZE_ALIGN. Add
!CONFIG_XDP_SOCKETS stubs for the new helpers and for two frame-size
helpers that had none, so drivers need no #ifdefs.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/xdp_sock_drv.h | 46 +++++++++++++++++++++++++++++++++++++-
net/xdp/xsk_buff_pool.c | 2 +-
2 files changed, 46 insertions(+), 2 deletions(-)
diff --git a/include/net/xdp_sock_drv.h b/include/net/xdp_sock_drv.h
index b9288f5dd48b..28461e675b8e 100644
--- a/include/net/xdp_sock_drv.h
+++ b/include/net/xdp_sock_drv.h
@@ -6,6 +6,7 @@
#ifndef _LINUX_XDP_SOCK_DRV_H
#define _LINUX_XDP_SOCK_DRV_H
+#include <net/netdev_rx_queue.h>
#include <net/xdp_sock.h>
#include <net/xsk_buff_pool.h>
@@ -13,6 +14,7 @@ struct netlink_ext_ack;
#define XDP_UMEM_MIN_CHUNK_SHIFT 11
#define XDP_UMEM_MIN_CHUNK_SIZE (1 << XDP_UMEM_MIN_CHUNK_SHIFT)
+#define XSK_RX_FRAME_SIZE_ALIGN 128
#define NETDEV_XDP_ACT_XSK (NETDEV_XDP_ACT_BASIC | \
NETDEV_XDP_ACT_REDIRECT | \
@@ -34,12 +36,33 @@ int xsk_pool_setup_page_pool(struct net_device *dev, struct xsk_buff_pool *pool,
u16 queue_id, struct netlink_ext_ack *extack);
struct xsk_buff_pool *xsk_get_pool_from_qid(struct net_device *dev,
u16 queue_id);
+
+/* Return the XSK pool only while it is the installed RX memory provider. */
+static inline struct xsk_buff_pool *
+xsk_get_pool_from_rxq(struct net_device *dev, u16 queue_id)
+{
+ struct netdev_rx_queue *rxq;
+ struct xsk_buff_pool *pool;
+
+ if (queue_id >= dev->real_num_rx_queues)
+ return NULL;
+
+ rxq = __netif_get_rx_queue(dev, queue_id);
+ pool = rxq->pool;
+ return pool == rxq->mp_params.mp_priv ? pool : NULL;
+}
+
void xsk_set_rx_need_wakeup(struct xsk_buff_pool *pool);
void xsk_set_tx_need_wakeup(struct xsk_buff_pool *pool);
void xsk_clear_rx_need_wakeup(struct xsk_buff_pool *pool);
void xsk_clear_tx_need_wakeup(struct xsk_buff_pool *pool);
bool xsk_uses_need_wakeup(struct xsk_buff_pool *pool);
+static inline bool xsk_pool_uses_sg(struct xsk_buff_pool *pool)
+{
+ return pool->umem->flags & XDP_UMEM_SG_FLAG;
+}
+
static inline u32 xsk_pool_get_headroom(struct xsk_buff_pool *pool)
{
return XDP_PACKET_HEADROOM + pool->headroom;
@@ -73,7 +96,7 @@ static inline u32 xsk_pool_get_rx_frame_size(struct xsk_buff_pool *pool)
mbuf = pool->dev && (umem->flags & XDP_UMEM_SG_FLAG);
frame_size -= xsk_pool_get_tailroom(mbuf);
- return ALIGN_DOWN(frame_size, 128);
+ return ALIGN_DOWN(frame_size, XSK_RX_FRAME_SIZE_ALIGN);
}
static inline u32 xsk_pool_get_rx_frag_step(struct xsk_buff_pool *pool)
@@ -388,6 +411,12 @@ xsk_get_pool_from_qid(struct net_device *dev, u16 queue_id)
return NULL;
}
+static inline struct xsk_buff_pool *
+xsk_get_pool_from_rxq(struct net_device *dev, u16 queue_id)
+{
+ return NULL;
+}
+
static inline void xsk_set_rx_need_wakeup(struct xsk_buff_pool *pool)
{
}
@@ -409,16 +438,31 @@ static inline bool xsk_uses_need_wakeup(struct xsk_buff_pool *pool)
return false;
}
+static inline bool xsk_pool_uses_sg(struct xsk_buff_pool *pool)
+{
+ return false;
+}
+
static inline u32 xsk_pool_get_headroom(struct xsk_buff_pool *pool)
{
return 0;
}
+static inline u32 xsk_pool_get_tailroom(bool mbuf)
+{
+ return 0;
+}
+
static inline u32 xsk_pool_get_chunk_size(struct xsk_buff_pool *pool)
{
return 0;
}
+static inline u32 __xsk_pool_get_rx_frame_size(struct xsk_buff_pool *pool)
+{
+ return 0;
+}
+
static inline u32 xsk_pool_get_rx_frame_size(struct xsk_buff_pool *pool)
{
return 0;
diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
index e25347f8c208..9c32fb165533 100644
--- a/net/xdp/xsk_buff_pool.c
+++ b/net/xdp/xsk_buff_pool.c
@@ -261,7 +261,7 @@ int xp_assign_dev(struct xsk_buff_pool *pool,
*/
frame_size = __xsk_pool_get_rx_frame_size(pool) -
xsk_pool_get_tailroom(mbuf);
- frame_size = ALIGN_DOWN(frame_size, 128);
+ frame_size = ALIGN_DOWN(frame_size, XSK_RX_FRAME_SIZE_ALIGN);
if (needed > frame_size * segs) {
err = -EINVAL;
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 11/15] xdp: Copy provider buffers on pass and redirect
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (9 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 10/15] xsk: Add RX helpers for page-pool drivers Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 12/15] xsk: Receive provider UMEM without copying Björn Töpel
` (3 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
A provider buffer must go back to its page pool from the NAPI poll
that allocated it. Userspace gets UMEM ownership back through the
FILL ring, without a reference count. An skb or xdp_frame can
outlive the poll, so it cannot hold a provider buffer.
Copy the packet, fragments included, into kernel pages on XDP_PASS and
on redirects to targets other than XSKMAP. On redirect, split the copy
from the release: return the provider buffer only after the target
accepts the copy. The classic zero-copy redirect path is unchanged.
The copy and the BPF fragment helpers get a fragment's address from
the new xdp_frag_address(). It works for pages and for provider
net_iovs; skb_frag_address() works only for pages. The next patch
handles delivery to an XSKMAP socket.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/xdp.h | 8 +++-
net/core/dev.h | 5 +++
net/core/filter.c | 28 +++++++++++---
net/core/xdp.c | 98 ++++++++++++++++++++++++++++++++++++++++++++---
4 files changed, 126 insertions(+), 13 deletions(-)
diff --git a/include/net/xdp.h b/include/net/xdp.h
index 80931490ca3a..4d7040debcca 100644
--- a/include/net/xdp.h
+++ b/include/net/xdp.h
@@ -420,11 +420,17 @@ xdp_update_skb_frags_info(struct sk_buff *skb, u8 nr_frags,
skb->unreadable |= !!(xdp_flags & XDP_FLAGS_FRAGS_UNREADABLE);
}
+/* Page and readable provider fragments alike resolve through their netmem. */
+static inline void *xdp_frag_address(const skb_frag_t *frag)
+{
+ return netmem_address(skb_frag_netmem(frag)) + skb_frag_off(frag);
+}
+
/* Avoids inlining WARN macro in fast-path */
void xdp_warn(const char *msg, const char *func, const int line);
#define XDP_WARN(msg) xdp_warn(msg, __func__, __LINE__)
-struct sk_buff *xdp_build_skb_from_buff(const struct xdp_buff *xdp);
+struct sk_buff *xdp_build_skb_from_buff(struct xdp_buff *xdp);
struct sk_buff *xdp_build_skb_from_zc(struct xdp_buff *xdp);
struct xdp_frame *xdp_convert_zc_to_xdp_frame(struct xdp_buff *xdp);
struct sk_buff *__xdp_build_skb_from_frame(struct xdp_frame *xdpf,
diff --git a/net/core/dev.h b/net/core/dev.h
index 4793a994f7f1..6598a4c5cfec 100644
--- a/net/core/dev.h
+++ b/net/core/dev.h
@@ -13,6 +13,11 @@ struct netlink_ext_ack;
struct netdev_queue_config;
struct cpumask;
struct pp_memory_provider_params;
+struct xdp_buff;
+struct xdp_frame;
+
+struct xdp_frame *xdp_copy_zc_to_xdp_frame(struct xdp_buff *xdp);
+void xdp_release_zc_buff(struct xdp_buff *xdp);
/* Random bits of netdevice that don't need to be exposed */
#define FLOW_LIMIT_HISTORY (1 << 7) /* must be ^2 and !overflow buckets */
diff --git a/net/core/filter.c b/net/core/filter.c
index 70dc621672f2..23b3923c0869 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -4221,7 +4221,7 @@ void bpf_xdp_copy_buf(struct xdp_buff *xdp, unsigned long off,
break;
ptr_off += ptr_len;
- ptr_buf = skb_frag_address(next_frag);
+ ptr_buf = xdp_frag_address(next_frag);
ptr_len = skb_frag_size(next_frag);
next_frag++;
}
@@ -4249,7 +4249,7 @@ void *bpf_xdp_pointer(struct xdp_buff *xdp, u32 offset, u32 len)
u32 frag_size = skb_frag_size(&sinfo->frags[i]);
if (offset < frag_size) {
- addr = skb_frag_address(&sinfo->frags[i]);
+ addr = xdp_frag_address(&sinfo->frags[i]);
size = frag_size;
break;
}
@@ -4339,7 +4339,7 @@ static int bpf_xdp_frags_increase_tail(struct xdp_buff *xdp, int offset)
if (unlikely(offset > tailroom))
return -EINVAL;
- memset(skb_frag_address(frag) + skb_frag_size(frag), 0, offset);
+ memset(xdp_frag_address(frag) + skb_frag_size(frag), 0, offset);
skb_frag_size_add(frag, offset);
sinfo->xdp_frags_size += offset;
if (rxq->mem.type == MEM_TYPE_XSK_BUFF_POOL)
@@ -4684,12 +4684,28 @@ int xdp_do_redirect(struct net_device *dev, struct xdp_buff *xdp,
{
struct bpf_redirect_info *ri = bpf_net_ctx_get_ri();
enum bpf_map_type map_type = ri->map_type;
+ struct xdp_frame *xdpf;
+ int err;
if (map_type == BPF_MAP_TYPE_XSKMAP)
return __xdp_do_redirect_xsk(ri, dev, xdp, xdp_prog);
- return __xdp_do_redirect_frame(ri, dev, xdp_convert_buff_to_frame(xdp),
- xdp_prog);
+ if (!xdp_buff_has_netmem(xdp))
+ return __xdp_do_redirect_frame(ri, dev,
+ xdp_convert_buff_to_frame(xdp),
+ xdp_prog);
+
+ /* Return provider buffers only after the target accepts the copy. */
+ xdpf = xdp_copy_zc_to_xdp_frame(xdp);
+ err = __xdp_do_redirect_frame(ri, dev, xdpf, xdp_prog);
+ if (err) {
+ if (xdpf)
+ xdp_return_frame_rx_napi(xdpf);
+ return err;
+ }
+
+ xdp_release_zc_buff(xdp);
+ return 0;
}
EXPORT_SYMBOL_GPL(xdp_do_redirect);
@@ -12677,7 +12693,7 @@ __bpf_kfunc int bpf_xdp_pull_data(struct xdp_md *x, u32 len)
skb_frag_t *frag = &sinfo->frags[i];
u32 shrink = min_t(u32, delta, skb_frag_size(frag));
- memcpy(xdp->data_end, skb_frag_address(frag), shrink);
+ memcpy(xdp->data_end, xdp_frag_address(frag), shrink);
xdp->data_end += shrink;
sinfo->xdp_frags_size -= shrink;
diff --git a/net/core/xdp.c b/net/core/xdp.c
index 386240bd24c9..b7f16f44dac2 100644
--- a/net/core/xdp.c
+++ b/net/core/xdp.c
@@ -23,6 +23,8 @@
#include <trace/events/xdp.h>
#include <net/xdp_sock_drv.h>
+#include "dev.h"
+
#define REG_STATE_NEW 0x0
#define REG_STATE_REGISTERED 0x1
#define REG_STATE_UNREGISTERED 0x2
@@ -559,7 +561,8 @@ void xdp_return_buff(struct xdp_buff *xdp)
xdp->rxq->mem.type, true, xdp);
out:
- __xdp_return(virt_to_netmem(xdp->data), xdp->rxq->mem.type, true, xdp);
+ __xdp_return(xdp_buff_get_netmem(xdp), xdp->rxq->mem.type, true,
+ xdp);
}
EXPORT_SYMBOL_GPL(xdp_return_buff);
@@ -573,7 +576,56 @@ void xdp_attachment_setup(struct xdp_attachment_info *info,
}
EXPORT_SYMBOL_GPL(xdp_attachment_setup);
-struct xdp_frame *xdp_convert_zc_to_xdp_frame(struct xdp_buff *xdp)
+static bool xdp_copy_frags_to_frame(struct xdp_frame *xdpf,
+ const struct xdp_buff *xdp)
+{
+ const struct skb_shared_info *xinfo;
+ struct skb_shared_info *sinfo;
+ u32 i, nr_frags;
+
+ xinfo = xdp_get_shared_info_from_buff(xdp);
+ nr_frags = xinfo->nr_frags;
+
+ sinfo = xdp_get_shared_info_from_frame(xdpf);
+ memset(sinfo, 0, sizeof(*sinfo));
+
+ for (i = 0; i < nr_frags; i++) {
+ const skb_frag_t *frag = &xinfo->frags[i];
+ u32 len = skb_frag_size(frag);
+ struct page *page;
+
+ page = dev_alloc_page();
+ if (!page)
+ goto err;
+
+ memcpy(page_address(page), xdp_frag_address(frag), len);
+ __skb_fill_page_desc_noacc(sinfo, i, page, 0, len);
+ if (page_is_pfmemalloc(page))
+ xdpf->flags |= XDP_FLAGS_FRAGS_PF_MEMALLOC;
+ }
+
+ sinfo->nr_frags = nr_frags;
+ sinfo->xdp_frags_size = xinfo->xdp_frags_size;
+ sinfo->xdp_frags_truesize = nr_frags * PAGE_SIZE;
+ return true;
+
+err:
+ while (i--)
+ put_page(skb_frag_page(&sinfo->frags[i]));
+
+ return false;
+}
+
+/**
+ * xdp_copy_zc_to_xdp_frame - copy a zero-copy buff into a page-backed frame
+ * @xdp: zero-copy &xdp_buff to copy
+ *
+ * The source keeps its buffers, so a caller that cannot hand the frame on
+ * must free the frame and may still return the source to its owner.
+ *
+ * Return: new frame on success, %NULL on failure.
+ */
+struct xdp_frame *xdp_copy_zc_to_xdp_frame(struct xdp_buff *xdp)
{
unsigned int metasize, totsize;
void *addr, *data_to_copy;
@@ -607,7 +659,34 @@ struct xdp_frame *xdp_convert_zc_to_xdp_frame(struct xdp_buff *xdp)
xdpf->frame_sz = PAGE_SIZE;
xdpf->mem_type = MEM_TYPE_PAGE_ORDER0;
- xsk_buff_free(xdp);
+ /* Classic zero-copy frames keep only the linear part. */
+ if (xdp_buff_has_netmem(xdp) && xdp_buff_has_frags(xdp)) {
+ xdpf->flags = XDP_FLAGS_HAS_FRAGS;
+ if (!xdp_copy_frags_to_frame(xdpf, xdp)) {
+ put_page(page);
+ return NULL;
+ }
+ }
+
+ return xdpf;
+}
+
+void xdp_release_zc_buff(struct xdp_buff *xdp)
+{
+ if (xdp->rxq->mem.type == MEM_TYPE_XSK_BUFF_POOL)
+ xsk_buff_free(xdp);
+ else
+ xdp_return_buff(xdp);
+}
+
+struct xdp_frame *xdp_convert_zc_to_xdp_frame(struct xdp_buff *xdp)
+{
+ struct xdp_frame *xdpf;
+
+ xdpf = xdp_copy_zc_to_xdp_frame(xdp);
+ if (xdpf)
+ xdp_release_zc_buff(xdp);
+
return xdpf;
}
EXPORT_SYMBOL_GPL(xdp_convert_zc_to_xdp_frame);
@@ -627,10 +706,11 @@ EXPORT_SYMBOL_GPL(xdp_warn);
* &xdp_buff: allocate an skb head from the NAPI percpu cache, initialize
* skb data pointers and offsets, set the recycle bit if the buff is
* PP-backed, Rx queue index, protocol and update frags info.
+ * Provider-backed netmem is copied into kernel memory and released.
*
* Return: new &sk_buff on success, %NULL on error.
*/
-struct sk_buff *xdp_build_skb_from_buff(const struct xdp_buff *xdp)
+struct sk_buff *xdp_build_skb_from_buff(struct xdp_buff *xdp)
{
const struct xdp_rxq_info *rxq = xdp->rxq;
const struct skb_shared_info *sinfo;
@@ -638,6 +718,9 @@ struct sk_buff *xdp_build_skb_from_buff(const struct xdp_buff *xdp)
u32 nr_frags = 0;
int metalen;
+ if (xdp_buff_has_netmem(xdp))
+ return xdp_build_skb_from_zc(xdp);
+
if (unlikely(xdp_buff_has_frags(xdp))) {
sinfo = xdp_get_shared_info_from_buff(xdp);
nr_frags = sinfo->nr_frags;
@@ -708,7 +791,7 @@ static noinline bool xdp_copy_frags_from_zc(struct sk_buff *skb,
return false;
}
- memcpy(page_address(page) + offset, skb_frag_address(frag),
+ memcpy(page_address(page) + offset, xdp_frag_address(frag),
len);
__skb_fill_page_desc_noacc(sinfo, i, page, offset, len);
@@ -787,7 +870,10 @@ struct sk_buff *xdp_build_skb_from_zc(struct xdp_buff *xdp)
goto out;
}
- xsk_buff_free(xdp);
+ if (xdp->rxq->mem.type == MEM_TYPE_XSK_BUFF_POOL)
+ xsk_buff_free(xdp);
+ else
+ xdp_return_buff(xdp);
skb->protocol = eth_type_trans(skb, rxq->dev);
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 12/15] xsk: Receive provider UMEM without copying
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (10 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 11/15] xdp: Copy provider buffers on pass and redirect Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 13/15] eth: fbnic: Support AF_XDP zero-copy receive Björn Töpel
` (2 subsequent siblings)
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
When all buffers of a packet come from the provider of the target
socket, put their UMEM addresses straight on the socket's RX ring.
For a single-buffer packet, xsk_rcv() checks the head buffer. For a
multi-buffer packet, every fragment is checked. A socket without
XDP_USE_SG drops multi-buffer packets and counts them in rx_dropped.
If any fragment comes from elsewhere, the packet is copied as
before. Fragment descriptors carry the offset that the device used.
The buffers leave page_pool ownership in batches when the socket is
flushed. A batch is split where a queue restart changed the page
pool.
The copy path now reads fragments with xdp_frag_address().
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
include/net/xdp_sock.h | 10 +++
net/xdp/xsk.c | 191 ++++++++++++++++++++++++++++++++++++++++-
2 files changed, 198 insertions(+), 3 deletions(-)
diff --git a/include/net/xdp_sock.h b/include/net/xdp_sock.h
index 6e70b320b399..92a7cf49f5e3 100644
--- a/include/net/xdp_sock.h
+++ b/include/net/xdp_sock.h
@@ -12,10 +12,18 @@
#include <linux/mutex.h>
#include <linux/spinlock.h>
#include <linux/mm.h>
+#include <net/netmem.h>
#include <net/sock.h>
#define XDP_UMEM_SG_FLAG BIT(3)
+/* Bound both one maximum-SG packet and single-buffer release batching. */
+#if MAX_SKB_FRAGS < 31
+#define XSK_PP_RELEASE_BATCH 32
+#else
+#define XSK_PP_RELEASE_BATCH (MAX_SKB_FRAGS + 1)
+#endif
+
struct net_device;
struct xsk_queue;
struct xdp_buff;
@@ -61,6 +69,8 @@ struct xdp_sock {
XSK_BOUND,
XSK_UNBOUND,
} state;
+ u32 pp_release_cnt;
+ netmem_ref pp_release[XSK_PP_RELEASE_BATCH];
struct xsk_queue *tx ____cacheline_aligned_in_smp;
struct list_head tx_list;
diff --git a/net/xdp/xsk.c b/net/xdp/xsk.c
index 90c98b42a18a..99bd2bc79467 100644
--- a/net/xdp/xsk.c
+++ b/net/xdp/xsk.c
@@ -30,6 +30,7 @@
#include <net/busy_poll.h>
#include <net/netdev_lock.h>
#include <net/netdev_rx_queue.h>
+#include <net/page_pool/memory_provider.h>
#include <net/xdp.h>
#include "../core/dev.h"
@@ -267,6 +268,157 @@ static int xsk_rcv_zc(struct xdp_sock *xs, struct xdp_buff *xdp, u32 len)
return err;
}
+static bool xsk_pp_can_xfer(struct xsk_buff_pool *pool, netmem_ref *netmems,
+ u32 count)
+{
+ u32 i;
+
+ if (WARN_ON_ONCE(!count || count > MAX_SKB_FRAGS + 1))
+ return false;
+
+ for (i = 0; i < count; i++)
+ if (!xp_netmem_is_from_pool(netmems[i], pool))
+ return false;
+
+ return true;
+}
+
+static void xsk_pp_release_to_user_bulk(netmem_ref *netmems, u32 count)
+{
+ struct page_pool *pp, *next;
+ u32 first = 0, i;
+
+ if (WARN_ON_ONCE(!count || count > XSK_PP_RELEASE_BATCH))
+ return;
+
+ /* One poll of one queue fills the batch, so it normally shares a
+ * page pool. Split it where a queue replacement changed the pool.
+ */
+ pp = netmem_get_pp(netmems[0]);
+ for (i = 1; i < count; i++) {
+ next = netmem_get_pp(netmems[i]);
+ if (next == pp)
+ continue;
+
+ net_mp_release_page_pool_bulk(pp, netmems + first, i - first);
+ pp = next;
+ first = i;
+ }
+ net_mp_release_page_pool_bulk(pp, netmems + first, count - first);
+}
+
+static void xsk_pp_flush(struct xdp_sock *xs)
+{
+ u32 count = xs->pp_release_cnt;
+
+ if (!count)
+ return;
+
+ /* Provider descriptors are the tail of the unpublished RX entries. */
+ xsk_pp_release_to_user_bulk(xs->pp_release, count);
+ xs->pp_release_cnt = 0;
+}
+
+static void xsk_pp_release(struct xdp_sock *xs, netmem_ref netmem)
+{
+ xs->pp_release[xs->pp_release_cnt++] = netmem;
+}
+
+static void xsk_rcv_pp_zc_desc(struct xdp_sock *xs, netmem_ref netmem,
+ void *data, u32 len, u32 flags)
+{
+ __xskq_prod_reserve_desc(xs->rx, data - xs->pool->addrs, len, flags);
+ xsk_pp_release(xs, netmem);
+}
+
+static __always_inline int xsk_rcv_pp_zc_one(struct xdp_sock *xs,
+ struct xdp_buff *xdp, u32 len)
+{
+ netmem_ref netmem = xdp_buff_get_netmem(xdp);
+
+ if (xskq_prod_nb_free(xs->rx, 1) < 1) {
+ xs->rx_queue_full++;
+ return -ENOBUFS;
+ }
+ if (unlikely(xs->pp_release_cnt == XSK_PP_RELEASE_BATCH)) {
+ xsk_pp_flush(xs);
+ xskq_prod_submit(xs->rx);
+ }
+
+ __xskq_prod_reserve_desc(xs->rx, xdp->data - xs->pool->addrs, len, 0);
+ xsk_pp_release(xs, netmem);
+ return 0;
+}
+
+static noinline int xsk_rcv_pp_zc_sg(struct xdp_sock *xs,
+ struct xdp_buff *xdp, u32 nr_frags)
+{
+ netmem_ref netmems[MAX_SKB_FRAGS + 1];
+ struct skb_shared_info *sinfo;
+ u32 num_desc;
+ u32 flags;
+ u32 len;
+ u32 i;
+
+ BUILD_BUG_ON(MAX_SKB_FRAGS + 1 > XSK_PP_RELEASE_BATCH);
+
+ if (xs->pp_release_cnt) {
+ xsk_pp_flush(xs);
+ xskq_prod_submit(xs->rx);
+ }
+
+ sinfo = xdp_get_shared_info_from_buff(xdp);
+ num_desc = nr_frags + 1;
+ netmems[0] = xdp_buff_get_netmem(xdp);
+ for (i = 0; i < nr_frags; i++)
+ netmems[i + 1] = skb_frag_netmem(&sinfo->frags[i]);
+
+ if (xskq_prod_nb_free(xs->rx, num_desc) < num_desc) {
+ xs->rx_queue_full++;
+ return -ENOBUFS;
+ }
+ if (!xsk_pp_can_xfer(xs->pool, netmems, num_desc))
+ return -EXDEV;
+
+ len = xdp->data_end - xdp->data;
+ flags = XDP_PKT_CONTD;
+ xsk_rcv_pp_zc_desc(xs, netmems[0], xdp->data, len, flags);
+
+ for (i = 0; i < nr_frags; i++) {
+ const skb_frag_t *frag = &sinfo->frags[i];
+ void *data = xdp_frag_address(frag);
+
+ if (i == nr_frags - 1)
+ flags = 0;
+
+ xsk_rcv_pp_zc_desc(xs, netmems[i + 1], data,
+ skb_frag_size(frag),
+ flags);
+ }
+ xsk_pp_flush(xs);
+ xskq_prod_submit(xs->rx);
+
+ return 0;
+}
+
+static int xsk_rcv_pp_zc(struct xdp_sock *xs, struct xdp_buff *xdp, u32 len)
+{
+ u32 nr_frags;
+
+ if (likely(!xdp_buff_has_frags(xdp)))
+ return xsk_rcv_pp_zc_one(xs, xdp, len);
+ if (unlikely(!xs->sg)) {
+ xs->rx_dropped++;
+ return -ENOSPC;
+ }
+
+ nr_frags = READ_ONCE(xdp_get_shared_info_from_buff(xdp)->nr_frags);
+ if (unlikely(!nr_frags || nr_frags > MAX_SKB_FRAGS))
+ return -EINVAL;
+
+ return xsk_rcv_pp_zc_sg(xs, xdp, nr_frags);
+}
+
static void *xsk_copy_xdp_start(struct xdp_buff *from)
{
if (unlikely(xdp_data_meta_unsupported(from)))
@@ -289,7 +441,7 @@ static u32 xsk_copy_xdp(void *to, void **from, u32 to_len,
return copied;
if (*from_len == copy_len) {
- *from = skb_frag_address(*frag);
+ *from = xdp_frag_address(*frag);
*from_len = skb_frag_size((*frag)++);
} else {
*from += copy_len;
@@ -461,12 +613,18 @@ static int xsk_rcv_check(struct xdp_sock *xs, struct xdp_buff *xdp, u32 len)
return 0;
}
-static void xsk_flush(struct xdp_sock *xs)
+static void __xsk_flush(struct xdp_sock *xs)
{
+ xsk_pp_flush(xs);
xskq_prod_submit(xs->rx);
/* Provider pools publish FILL consumption under the provider lock. */
if (!READ_ONCE(xs->pool->pp))
__xskq_cons_release(xs->pool->fq);
+}
+
+static void xsk_flush(struct xdp_sock *xs)
+{
+ __xsk_flush(xs);
sock_def_readable(&xs->sk);
}
@@ -485,8 +643,9 @@ int xsk_generic_rcv(struct xdp_sock *xs, struct xdp_buff *xdp)
return -EOPNOTSUPP;
}
err = __xsk_rcv(xs, xdp, len);
- xsk_flush(xs);
+ __xsk_flush(xs);
spin_unlock_bh(&xs->pool->rx_lock);
+ sock_def_readable(&xs->sk);
return err;
}
@@ -520,9 +679,31 @@ static int xsk_rcv(struct xdp_sock *xs, struct xdp_buff *xdp)
return err;
if (xdp->rxq->mem.type == MEM_TYPE_XSK_BUFF_POOL) {
+ if (unlikely(xs->pp_release_cnt)) {
+ xsk_pp_flush(xs);
+ xskq_prod_submit(xs->rx);
+ }
len = xdp->data_end - xdp->data;
return xsk_rcv_zc(xs, xdp, len);
}
+ if (xdp->rxq->mem.type == MEM_TYPE_PAGE_POOL &&
+ xdp_buff_has_netmem(xdp) &&
+ xp_netmem_is_from_pool(xdp_buff_get_netmem(xdp), xs->pool)) {
+ err = xsk_rcv_pp_zc(xs, xdp, len);
+ if (err != -EXDEV)
+ return err;
+
+ /* The redirect originates from this pool's registered RXQ, so
+ * the copy fallback produces into the socket RX ring from the
+ * same poll context as zero-copy delivery.
+ */
+ xsk_pp_flush(xs);
+ xskq_prod_submit(xs->rx);
+ err = xsk_rcv_copy(xs, xdp, len);
+ if (!err)
+ xdp_return_buff(xdp);
+ return err;
+ }
/* The socket RX ring has a single producer, the queue's poll context.
* Reject synthetic RX queues before their remote context produces
@@ -532,6 +713,10 @@ static int xsk_rcv(struct xdp_sock *xs, struct xdp_buff *xdp)
!xdp_rxq_info_is_reg(xdp->rxq)))
return -EINVAL;
+ if (unlikely(xs->pp_release_cnt)) {
+ xsk_pp_flush(xs);
+ xskq_prod_submit(xs->rx);
+ }
err = xsk_rcv_copy(xs, xdp, len);
if (!err)
xdp_return_buff(xdp);
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 13/15] eth: fbnic: Support AF_XDP zero-copy receive
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (11 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 12/15] xsk: Receive provider UMEM without copying Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 14/15] eth: fbnic: Support AF_XDP zero-copy transmit Björn Töpel
2026-10-02 19:00 ` [RFC net-next 15/15] Documentation: xsk: Document page-pool zero copy Björn Töpel
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
fbnic gets RX buffers from page_pool. Use the XSK buffer pool as the
RX queue's page-pool memory provider, so allocation, DMA sync,
recycling and refill stay in page_pool. Only aligned 4 KiB UMEM
chunks are supported.
Each UMEM frame on the header queue must hold one packet. Raise the
header-split threshold to at least 3072 bytes, more than half of a
4 KiB page, so the device finishes the frame after one header. Cap
it to the space left in the frame after the headroom. An MTU above
the threshold needs a socket with XDP_USE_SG and an XDP program with
frags support. With XDP_USE_SG, the payload queue also uses UMEM
frames, and the device uses each frame for one descriptor only.
Without it, the payload queue uses a normal page pool. If the device
does not finish a provider frame after one descriptor, drop the
packet.
The BDQ refill loop now alternates header and payload buffers on all
queues, so two rings that share one finite provider stay balanced.
It keeps NAPI scheduled while the provider still has buffers.
Provider frames are never split, so they get no fragment bias. A
redirected frame can be recycled before its old BDQ slot is cleaned,
and a stale bias would corrupt the new allocation.
The headroom comes from the queue configuration. It must be at least
XDP_PACKET_HEADROOM, a multiple of 128 bytes, and fit the hardware
field; otherwise queue validation fails. MTU, ring parameter and XDP
program changes that do not fit an active provider are rejected.
fbnic did not support XDP_REDIRECT. Add it, with xdp_do_flush(), and
advertise it. This also enables copy-mode AF_XDP and the other
redirect targets on fbnic. Do not advertise AF_XDP zero copy until
TX support is added.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
.../net/ethernet/meta/fbnic/fbnic_ethtool.c | 5 +
.../net/ethernet/meta/fbnic/fbnic_netdev.c | 140 ++++++-
.../net/ethernet/meta/fbnic/fbnic_netdev.h | 4 +
drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 350 +++++++++++++-----
drivers/net/ethernet/meta/fbnic/fbnic_txrx.h | 6 +
5 files changed, 402 insertions(+), 103 deletions(-)
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c
index e774740fcb5d..188abe0ead14 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c
@@ -385,6 +385,11 @@ fbnic_set_ringparam(struct net_device *netdev, struct ethtool_ringparam *ring,
return -EINVAL;
}
+ err = fbnic_xsk_validate(fbn, fbn->xdp_prog, netdev->mtu,
+ kernel_ring->hds_thresh, extack);
+ if (err)
+ return err;
+
if (!netif_running(netdev)) {
fbnic_set_rings(fbn, ring, kernel_ring);
return 0;
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c b/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c
index 8a6703afca38..038e91ef14e4 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c
@@ -5,6 +5,8 @@
#include <linux/ipv6.h>
#include <linux/types.h>
#include <net/netdev_queues.h>
+#include <net/netdev_rx_queue.h>
+#include <net/xdp_sock_drv.h>
#include "fbnic.h"
#include "fbnic_netdev.h"
@@ -276,6 +278,8 @@ static int fbnic_set_mac(struct net_device *netdev, void *p)
static int fbnic_change_mtu(struct net_device *dev, int new_mtu)
{
struct fbnic_net *fbn = netdev_priv(dev);
+ struct netlink_ext_ack extack = {};
+ int err;
if (fbnic_check_split_frames(fbn->xdp_prog, new_mtu, fbn->hds_thresh)) {
dev_err(&dev->dev,
@@ -285,6 +289,15 @@ static int fbnic_change_mtu(struct net_device *dev, int new_mtu)
return -EINVAL;
}
+ err = fbnic_xsk_validate(fbn, fbn->xdp_prog, new_mtu,
+ fbn->hds_thresh, &extack);
+ if (err) {
+ if (extack._msg)
+ netdev_err(dev, "%s\n", extack._msg);
+
+ return err;
+ }
+
WRITE_ONCE(dev->mtu, new_mtu);
return 0;
@@ -534,26 +547,123 @@ bool fbnic_check_split_frames(struct bpf_prog *prog, unsigned int mtu,
return mtu + ETH_HLEN > hds_thresh;
}
+u32 fbnic_xsk_hds_thresh(u32 headroom, u32 hds_thresh)
+{
+ u32 max_hds;
+
+ /* Reserve more than half a device page for each header so the device
+ * retires the provider frame after one descriptor.
+ */
+ static_assert(2 * FBNIC_XSK_HDS_THRESH > FBNIC_BD_PAGE_SIZE);
+
+ max_hds = FBNIC_BD_PAGE_SIZE - headroom - FBNIC_RX_TROOM -
+ FBNIC_RX_PAD;
+
+ return min(max_t(u32, hds_thresh, FBNIC_XSK_HDS_THRESH), max_hds);
+}
+
+static int fbnic_xsk_validate_frames(struct xsk_buff_pool *pool,
+ struct bpf_prog *prog,
+ unsigned int mtu, u32 hds_thresh,
+ struct netlink_ext_ack *extack)
+{
+ u32 xsk_hds_thresh;
+
+ xsk_hds_thresh = fbnic_xsk_hds_thresh(xsk_pool_get_headroom(pool),
+ hds_thresh);
+
+ if (mtu + ETH_HLEN <= xsk_hds_thresh)
+ return 0;
+
+ if (!xsk_pool_uses_sg(pool)) {
+ NL_SET_ERR_MSG_MOD(extack,
+ "AF_XDP zero-copy socket requires XDP_USE_SG");
+ return -EOPNOTSUPP;
+ }
+
+ if (fbnic_check_split_frames(prog, mtu, xsk_hds_thresh)) {
+ NL_SET_ERR_MSG_MOD(extack,
+ "AF_XDP zero-copy socket requires XDP frags");
+ return -EOPNOTSUPP;
+ }
+
+ return 0;
+}
+
+int fbnic_xsk_validate(struct fbnic_net *fbn, struct bpf_prog *prog,
+ unsigned int mtu, u32 hds_thresh,
+ struct netlink_ext_ack *extack)
+{
+ struct net_device *netdev = fbn->netdev;
+ struct xsk_buff_pool *pool;
+ unsigned int qid;
+ int err;
+
+ for (qid = 0; qid < fbn->num_rx_queues; qid++) {
+ pool = xsk_get_pool_from_rxq(netdev, qid);
+ if (!pool)
+ continue;
+
+ err = fbnic_xsk_validate_frames(pool, prog, mtu, hds_thresh,
+ extack);
+ if (err)
+ return err;
+ }
+
+ return 0;
+}
+
+static int fbnic_xsk_setup_pool(struct net_device *netdev,
+ struct xsk_buff_pool *pool, u16 queue_id,
+ struct netlink_ext_ack *extack)
+{
+ struct fbnic_net *fbn = netdev_priv(netdev);
+ int err;
+
+ if (queue_id >= fbn->num_rx_queues)
+ return -EINVAL;
+
+ if (!pool)
+ return xsk_pool_setup_page_pool(netdev, NULL, queue_id, extack);
+
+ err = fbnic_xsk_validate_frames(pool, fbn->xdp_prog, netdev->mtu,
+ fbn->hds_thresh, extack);
+ if (err)
+ return err;
+
+ return xsk_pool_setup_page_pool(netdev, pool, queue_id, extack);
+}
+
static int fbnic_bpf(struct net_device *netdev, struct netdev_bpf *bpf)
{
struct bpf_prog *prog = bpf->prog, *prev_prog;
struct fbnic_net *fbn = netdev_priv(netdev);
+ int err;
- if (bpf->command != XDP_SETUP_PROG)
+ switch (bpf->command) {
+ case XDP_SETUP_PROG:
+ if (fbnic_check_split_frames(prog, netdev->mtu,
+ fbn->hds_thresh)) {
+ NL_SET_ERR_MSG_MOD(bpf->extack,
+ "MTU too high, or HDS threshold is too low for single buffer XDP");
+ return -EOPNOTSUPP;
+ }
+
+ err = fbnic_xsk_validate(fbn, prog, netdev->mtu,
+ fbn->hds_thresh, bpf->extack);
+ if (err)
+ return err;
+
+ prev_prog = xchg(&fbn->xdp_prog, prog);
+ if (prev_prog)
+ bpf_prog_put(prev_prog);
+ return 0;
+ case XDP_SETUP_XSK_POOL:
+ return fbnic_xsk_setup_pool(netdev, bpf->xsk.pool,
+ bpf->xsk.queue_id, bpf->extack);
+ default:
return -EINVAL;
-
- if (fbnic_check_split_frames(prog, netdev->mtu,
- fbn->hds_thresh)) {
- NL_SET_ERR_MSG_MOD(bpf->extack,
- "MTU too high, or HDS threshold is too low for single buffer XDP");
- return -EOPNOTSUPP;
}
-
- prev_prog = xchg(&fbn->xdp_prog, prog);
- if (prev_prog)
- bpf_prog_put(prev_prog);
-
- return 0;
}
static const struct net_device_ops fbnic_netdev_ops = {
@@ -822,7 +932,9 @@ struct net_device *fbnic_netdev_alloc(struct fbnic_dev *fbd)
netdev->hw_enc_features |= netdev->features;
netdev->features |= NETIF_F_NTUPLE;
- netdev->xdp_features = NETDEV_XDP_ACT_BASIC | NETDEV_XDP_ACT_RX_SG;
+ netdev->xdp_features = NETDEV_XDP_ACT_BASIC |
+ NETDEV_XDP_ACT_REDIRECT |
+ NETDEV_XDP_ACT_RX_SG;
netdev->min_mtu = IPV6_MIN_MTU;
netdev->max_mtu = FBNIC_MAX_JUMBO_FRAME_SIZE - ETH_HLEN;
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_netdev.h b/drivers/net/ethernet/meta/fbnic/fbnic_netdev.h
index eded20b0e9e4..ba425302ccfd 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_netdev.h
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_netdev.h
@@ -116,4 +116,8 @@ int fbnic_phylink_init(struct net_device *netdev);
void fbnic_phylink_pmd_training_complete_notify(struct net_device *netdev);
bool fbnic_check_split_frames(struct bpf_prog *prog,
unsigned int mtu, u32 hds_threshold);
+int fbnic_xsk_validate(struct fbnic_net *fbn, struct bpf_prog *prog,
+ unsigned int mtu, u32 hds_thresh,
+ struct netlink_ext_ack *extack);
+u32 fbnic_xsk_hds_thresh(u32 headroom, u32 hds_thresh);
#endif /* _FBNIC_NETDEV_H_ */
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
index 0c7b1ea002d8..e6043607d4e2 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
@@ -12,6 +12,7 @@
#include <net/page_pool/helpers.h>
#include <net/tcp.h>
#include <net/xdp.h>
+#include <net/xdp_sock_drv.h>
#include "fbnic.h"
#include "fbnic_csr.h"
@@ -22,6 +23,7 @@ enum {
FBNIC_XDP_PASS = 0,
FBNIC_XDP_CONSUME,
FBNIC_XDP_TX,
+ FBNIC_XDP_REDIRECT,
FBNIC_XDP_LEN_ERR,
};
@@ -758,21 +760,29 @@ static void fbnic_page_pool_init(struct fbnic_ring *ring, unsigned int idx,
netmem_ref netmem)
{
struct fbnic_rx_buf *rx_buf = &ring->rx_buf[idx];
+ unsigned int pagecnt_bias;
- page_pool_fragment_netmem(netmem, FBNIC_PAGECNT_BIAS_MAX);
- rx_buf->pagecnt_bias = FBNIC_PAGECNT_BIAS_MAX;
+ /* Unsplittable provider frames can be recycled before the BDQ entry is
+ * cleaned.
+ */
+ if (page_pool_supports_frag(ring->page_pool)) {
+ pagecnt_bias = FBNIC_PAGECNT_BIAS_MAX;
+ page_pool_fragment_netmem(netmem, pagecnt_bias);
+ } else {
+ pagecnt_bias = 1;
+ }
+ rx_buf->pagecnt_bias = pagecnt_bias;
rx_buf->netmem = netmem;
}
-static struct page *
+static netmem_ref
fbnic_page_pool_get_head(struct fbnic_q_triad *qt, unsigned int idx)
{
struct fbnic_rx_buf *rx_buf = &qt->sub0.rx_buf[idx];
rx_buf->pagecnt_bias--;
- /* sub0 is always fed system pages, from the NAPI-level page_pool */
- return netmem_to_page(rx_buf->netmem);
+ return rx_buf->netmem;
}
static netmem_ref
@@ -791,7 +801,8 @@ static void fbnic_page_pool_drain(struct fbnic_ring *ring, unsigned int idx,
struct fbnic_rx_buf *rx_buf = &ring->rx_buf[idx];
netmem_ref netmem = rx_buf->netmem;
- if (!page_pool_unref_netmem(netmem, rx_buf->pagecnt_bias))
+ if (rx_buf->pagecnt_bias &&
+ !page_pool_unref_netmem(netmem, rx_buf->pagecnt_bias))
page_pool_put_unrefed_netmem(ring->page_pool, netmem, -1,
!!budget);
@@ -930,66 +941,121 @@ static void fbnic_bd_prep(struct fbnic_ring *bdq, u16 id, netmem_ref netmem)
} while (--i);
}
-static void fbnic_fill_bdq(struct fbnic_ring *bdq)
+static bool fbnic_fill_bdq_one(struct fbnic_ring *bdq, unsigned int *tail)
{
- unsigned int count = fbnic_desc_unused(bdq);
- unsigned int i = bdq->tail;
+ netmem_ref netmem;
- if (!count)
- return;
-
- do {
- netmem_ref netmem;
-
- netmem = page_pool_dev_alloc_netmems(bdq->page_pool);
- if (!netmem) {
- u64_stats_update_begin(&bdq->stats.syncp);
- bdq->stats.bdq.alloc_failed++;
- u64_stats_update_end(&bdq->stats.syncp);
-
- break;
- }
-
- fbnic_page_pool_init(bdq, i, netmem);
- fbnic_bd_prep(bdq, i, netmem);
-
- i++;
- i &= bdq->size_mask;
-
- count--;
- } while (count);
-
- if (bdq->tail != i) {
- bdq->tail = i;
-
- /* Force DMA writes to flush before writing to tail */
- dma_wmb();
-
- writel(i * fbnic_bd_page_count(bdq), bdq->doorbell);
+ netmem = page_pool_dev_alloc_netmems(bdq->page_pool);
+ if (!netmem) {
+ u64_stats_update_begin(&bdq->stats.syncp);
+ bdq->stats.bdq.alloc_failed++;
+ u64_stats_update_end(&bdq->stats.syncp);
+ return false;
}
+
+ fbnic_page_pool_init(bdq, *tail, netmem);
+ fbnic_bd_prep(bdq, *tail, netmem);
+ *tail = (*tail + 1) & bdq->size_mask;
+
+ return true;
}
-static unsigned int fbnic_hdr_pg_start(unsigned int pg_off)
+static void fbnic_fill_bdq_done(struct fbnic_ring *bdq, unsigned int tail)
+{
+ if (tail == bdq->tail)
+ return;
+
+ bdq->tail = tail;
+
+ /* Force DMA writes to flush before writing to tail */
+ dma_wmb();
+ writel(tail * fbnic_bd_page_count(bdq), bdq->doorbell);
+}
+
+static bool fbnic_fill_bdqs(struct fbnic_q_triad *qt)
+{
+ unsigned int hpq_count = fbnic_desc_unused(&qt->sub0);
+ unsigned int ppq_count = fbnic_desc_unused(&qt->sub1);
+ unsigned int hpq_tail = qt->sub0.tail;
+ unsigned int ppq_tail = qt->sub1.tail;
+ bool same_pool = qt->sub0.page_pool == qt->sub1.page_pool;
+ bool hpq_alloc = true;
+ bool ppq_alloc = true;
+ bool complete;
+
+ while ((hpq_alloc && hpq_count) || (ppq_alloc && ppq_count)) {
+ if (hpq_alloc && hpq_count) {
+ hpq_alloc = fbnic_fill_bdq_one(&qt->sub0, &hpq_tail);
+ hpq_count--;
+ }
+ if (ppq_alloc && ppq_count) {
+ ppq_alloc = fbnic_fill_bdq_one(&qt->sub1, &ppq_tail);
+ ppq_count--;
+ }
+ }
+
+ fbnic_fill_bdq_done(&qt->sub0, hpq_tail);
+ fbnic_fill_bdq_done(&qt->sub1, ppq_tail);
+
+ if (same_pool)
+ return page_pool_refill_done(qt->sub0.page_pool,
+ !fbnic_desc_unused(&qt->sub0) &&
+ !fbnic_desc_unused(&qt->sub1));
+
+ complete = page_pool_refill_done(qt->sub0.page_pool,
+ !fbnic_desc_unused(&qt->sub0));
+ complete &= page_pool_refill_done(qt->sub1.page_pool,
+ !fbnic_desc_unused(&qt->sub1));
+
+ return complete;
+}
+
+static unsigned int fbnic_hdr_pg_start(unsigned int pg_off,
+ unsigned int hroom)
{
/* The headroom of the first header may be larger than FBNIC_RX_HROOM
* due to alignment. So account for that by just making the page
* offset 0 if we are starting at the first header.
*/
- if (ALIGN(FBNIC_RX_HROOM, 128) > FBNIC_RX_HROOM &&
- pg_off == ALIGN(FBNIC_RX_HROOM, 128))
+ if (ALIGN(hroom, FBNIC_RX_HROOM_PAD) > hroom &&
+ pg_off == ALIGN(hroom, FBNIC_RX_HROOM_PAD))
return 0;
- return pg_off - FBNIC_RX_HROOM;
+ return pg_off - hroom;
}
-static unsigned int fbnic_hdr_pg_end(unsigned int pg_off, unsigned int len)
+static unsigned int fbnic_hdr_pg_end(unsigned int pg_off, unsigned int len,
+ unsigned int hroom)
{
/* Determine the end of the buffer by finding the start of the next
* and then subtracting the headroom from that frame.
*/
- pg_off += len + FBNIC_RX_TROOM + FBNIC_RX_HROOM;
+ pg_off += len + FBNIC_RX_TROOM + hroom;
- return ALIGN(pg_off, 128) - FBNIC_RX_HROOM;
+ return ALIGN(pg_off, FBNIC_RX_HROOM_PAD) - hroom;
+}
+
+static bool
+fbnic_netmem_page_finished(struct fbnic_napi_vector *nv,
+ struct fbnic_ring *bdq, unsigned int q_idx, u64 rcd)
+{
+ unsigned int idx = fbnic_rcd_bd_idx(bdq, rcd);
+ long bias = bdq->rx_buf[idx].pagecnt_bias;
+ bool page_finished;
+
+ page_finished = FIELD_GET(FBNIC_RCD_AL_PAGE_FIN, rcd);
+
+ /* Fragmentable pages start with a bias greater than one. */
+ if (bias > 1 || (bias == 1 && page_finished))
+ return true;
+
+ netdev_err_once(nv->napi.dev,
+ "Provider buffer not retired: queue %u id %u offset %llu len %llu\n",
+ q_idx, idx,
+ FIELD_GET(FBNIC_RCD_AL_BUFF_OFF_MASK, rcd),
+ FIELD_GET(FBNIC_RCD_AL_BUFF_LEN_MASK, rcd));
+
+ return false;
}
static void fbnic_pkt_prepare(struct fbnic_napi_vector *nv, u64 rcd,
@@ -1001,36 +1067,44 @@ static void fbnic_pkt_prepare(struct fbnic_napi_vector *nv, u64 rcd,
unsigned int hdr_pg_idx = fbnic_rcd_bd_idx(&qt->sub0, rcd);
unsigned int frame_sz, hdr_pg_start, hdr_pg_end, headroom;
unsigned char *hdr_start;
- struct page *page;
+ netmem_ref netmem;
/* data_hard_start should always be NULL when this is called */
WARN_ON_ONCE(pkt->buff.data_hard_start);
+ pkt->add_frag_failed = false;
+ if (unlikely(!fbnic_netmem_page_finished(nv, &qt->sub0,
+ qt->cmpl.q_idx, rcd))) {
+ pkt->add_frag_failed = true;
+ return;
+ }
- page = fbnic_page_pool_get_head(qt, hdr_pg_idx);
+ netmem = fbnic_page_pool_get_head(qt, hdr_pg_idx);
/* Short-cut the end calculation if we know page is fully consumed */
hdr_pg_end = FIELD_GET(FBNIC_RCD_AL_PAGE_FIN, rcd) ?
- FBNIC_BD_PAGE_SIZE : fbnic_hdr_pg_end(hdr_pg_off, len);
- hdr_pg_start = fbnic_hdr_pg_start(hdr_pg_off);
+ FBNIC_BD_PAGE_SIZE :
+ fbnic_hdr_pg_end(hdr_pg_off, len, qt->rx_hroom);
+ hdr_pg_start = fbnic_hdr_pg_start(hdr_pg_off, qt->rx_hroom);
+ hdr_start = netmem_address(netmem);
+ hdr_pg_start += qt->rx_frame_off;
headroom = hdr_pg_off - hdr_pg_start + FBNIC_RX_PAD;
frame_sz = hdr_pg_end - hdr_pg_start;
- xdp_init_buff(&pkt->buff, frame_sz, &qt->xdp_rxq);
+ xdp_init_buff_from_netmem(&pkt->buff, frame_sz, &qt->xdp_rxq,
+ netmem, &pkt->sinfo);
hdr_pg_start += fbnic_rcd_bd_page_offset(&qt->sub0, rcd);
/* Sync DMA buffer */
- dma_sync_single_range_for_cpu(nv->dev, page_pool_get_dma_addr(page),
- hdr_pg_start, frame_sz,
- DMA_BIDIRECTIONAL);
+ page_pool_dma_sync_netmem_for_cpu(qt->sub0.page_pool, netmem,
+ hdr_pg_start, frame_sz);
/* Build frame around buffer */
- hdr_start = page_address(page) + hdr_pg_start;
+ hdr_start += hdr_pg_start;
net_prefetch(pkt->buff.data);
xdp_prepare_buff(&pkt->buff, hdr_start, headroom,
len - FBNIC_RX_PAD, true);
pkt->hwtstamp = 0;
- pkt->add_frag_failed = false;
}
static void fbnic_add_rx_frag(struct fbnic_napi_vector *nv, u64 rcd,
@@ -1044,6 +1118,12 @@ static void fbnic_add_rx_frag(struct fbnic_napi_vector *nv, u64 rcd,
netmem_ref netmem;
bool added;
+ if (unlikely(!fbnic_netmem_page_finished(nv, &qt->sub1,
+ qt->cmpl.q_idx, rcd))) {
+ pkt->add_frag_failed = true;
+ return;
+ }
+
netmem = fbnic_page_pool_get_data(qt, pg_idx);
truesize = FIELD_GET(FBNIC_RCD_AL_PAGE_FIN, rcd) ?
@@ -1052,6 +1132,12 @@ static void fbnic_add_rx_frag(struct fbnic_napi_vector *nv, u64 rcd,
pg_off += fbnic_rcd_bd_page_offset(&qt->sub1, rcd);
+ /* The head buffer cannot safely hold fragment metadata. */
+ if (unlikely(pkt->add_frag_failed)) {
+ page_pool_put_full_netmem(qt->sub1.page_pool, netmem, true);
+ return;
+ }
+
/* Sync DMA buffer */
page_pool_dma_sync_netmem_for_cpu(qt->sub1.page_pool, netmem,
pg_off, truesize);
@@ -1071,8 +1157,6 @@ static void fbnic_add_rx_frag(struct fbnic_napi_vector *nv, u64 rcd,
static void fbnic_put_pkt_buff(struct fbnic_q_triad *qt,
struct fbnic_pkt_buff *pkt, int budget)
{
- struct page *page;
-
if (!pkt->buff.data_hard_start)
return;
@@ -1091,8 +1175,8 @@ static void fbnic_put_pkt_buff(struct fbnic_q_triad *qt,
}
}
- page = virt_to_page(pkt->buff.data_hard_start);
- page_pool_put_full_page(qt->sub0.page_pool, page, !!budget);
+ page_pool_put_full_netmem(qt->sub0.page_pool,
+ xdp_buff_get_netmem(&pkt->buff), !!budget);
}
static struct sk_buff *fbnic_build_skb(struct fbnic_napi_vector *nv,
@@ -1211,6 +1295,10 @@ static struct sk_buff *fbnic_run_xdp(struct fbnic_napi_vector *nv,
return fbnic_build_skb(nv, pkt);
case XDP_TX:
return ERR_PTR(fbnic_pkt_tx(nv, pkt));
+ case XDP_REDIRECT:
+ if (xdp_do_redirect(nv->napi.dev, &pkt->buff, xdp_prog))
+ break;
+ return ERR_PTR(-FBNIC_XDP_REDIRECT);
default:
bpf_warn_invalid_xdp_action(nv->napi.dev, xdp_prog, act);
fallthrough;
@@ -1280,6 +1368,8 @@ static int fbnic_clean_rcq(struct fbnic_napi_vector *nv,
struct fbnic_ring *rcq = &qt->cmpl;
struct fbnic_pkt_buff *pkt;
__le64 *raw_rcd, done;
+ bool redirect = false;
+ bool complete;
u32 head = rcq->head;
done = (head & (rcq->size_mask + 1)) ? cpu_to_le64(FBNIC_RCD_DONE) : 0;
@@ -1334,6 +1424,8 @@ static int fbnic_clean_rcq(struct fbnic_napi_vector *nv,
napi_gro_receive(&nv->napi, skb);
} else if (skb == ERR_PTR(-FBNIC_XDP_TX)) {
pkt_tail = nv->qt[0].sub1.tail;
+ } else if (skb == ERR_PTR(-FBNIC_XDP_REDIRECT)) {
+ redirect = true;
} else if (PTR_ERR(skb) == -FBNIC_XDP_CONSUME) {
fbnic_put_pkt_buff(qt, pkt, 1);
} else {
@@ -1377,15 +1469,17 @@ static int fbnic_clean_rcq(struct fbnic_napi_vector *nv,
if (pkt_tail >= 0)
fbnic_pkt_commit_tail(nv, pkt_tail);
+ if (redirect)
+ xdp_do_flush();
/* Unmap and free processed buffers */
if (head0 >= 0)
fbnic_clean_bdq(&qt->sub0, head0, budget);
- fbnic_fill_bdq(&qt->sub0);
if (head1 >= 0)
fbnic_clean_bdq(&qt->sub1, head1, budget);
- fbnic_fill_bdq(&qt->sub1);
+
+ complete = fbnic_fill_bdqs(qt);
/* Record the current head/tail of the queue */
if (rcq->head != head) {
@@ -1393,7 +1487,7 @@ static int fbnic_clean_rcq(struct fbnic_napi_vector *nv,
writel(head & rcq->size_mask, rcq->doorbell);
}
- return packets;
+ return complete ? packets : budget;
}
static void fbnic_nv_irq_disable(struct fbnic_napi_vector *nv)
@@ -1594,8 +1688,10 @@ void fbnic_free_napi_vectors(struct fbnic_net *fbn)
static int
fbnic_alloc_qt_page_pools(struct fbnic_net *fbn, struct fbnic_q_triad *qt,
- unsigned int rxq_idx, u32 rx_page_size)
+ unsigned int rxq_idx,
+ const struct netdev_queue_config *qcfg)
{
+ u32 rx_page_size = qcfg->rx_page_size;
struct page_pool_params pp_params = {
.order = 0,
.flags = PP_FLAG_DMA_MAP |
@@ -1609,7 +1705,12 @@ fbnic_alloc_qt_page_pools(struct fbnic_net *fbn, struct fbnic_q_triad *qt,
.netdev = fbn->netdev,
.queue_idx = rxq_idx,
};
+ struct xsk_buff_pool *xsk_pool = NULL;
struct page_pool *pp;
+ bool xsk_sg = false;
+
+ qt->rx_hroom = FBNIC_RX_HROOM;
+ qt->rx_frame_off = 0;
/* Page pool cannot exceed a size of 32768. This doesn't limit the
* pages on the ring but the number we can have cached waiting on
@@ -1623,6 +1724,21 @@ fbnic_alloc_qt_page_pools(struct fbnic_net *fbn, struct fbnic_q_triad *qt,
if (pp_params.pool_size > 32768)
pp_params.pool_size = 32768;
+ xsk_pool = xsk_get_pool_from_rxq(fbn->netdev, rxq_idx);
+ qt->xsk_pool = xsk_pool;
+ if (xsk_pool) {
+ xsk_sg = xsk_pool_uses_sg(xsk_pool);
+ pp_params.flags |= PP_FLAG_ALLOW_UNREADABLE_NETMEM;
+ }
+
+ /* A memory provider reserves headroom in front of the XDP headroom.
+ * The frame, and data_hard_start, begin after it.
+ */
+ if (qcfg->rx_headroom) {
+ qt->rx_hroom = qcfg->rx_headroom;
+ qt->rx_frame_off = qcfg->rx_headroom - XDP_PACKET_HEADROOM;
+ }
+
pp = page_pool_create(&pp_params);
if (IS_ERR(pp))
return PTR_ERR(pp);
@@ -1634,6 +1750,12 @@ fbnic_alloc_qt_page_pools(struct fbnic_net *fbn, struct fbnic_q_triad *qt,
pp_params.flags |= PP_FLAG_ALLOW_UNREADABLE_NETMEM;
pp_params.dma_dir = DMA_FROM_DEVICE;
+ pp = page_pool_create(&pp_params);
+ if (IS_ERR(pp))
+ goto err_destroy_sub0;
+ } else if (xsk_pool && !xsk_sg) {
+ pp_params.flags &= ~PP_FLAG_ALLOW_UNREADABLE_NETMEM;
+
pp = page_pool_create(&pp_params);
if (IS_ERR(pp))
goto err_destroy_sub0;
@@ -2058,19 +2180,20 @@ static int fbnic_alloc_tx_qt_resources(struct fbnic_net *fbn,
return err;
}
-static int fbnic_alloc_rx_qt_resources(struct fbnic_net *fbn,
- struct fbnic_napi_vector *nv,
- struct fbnic_q_triad *qt,
- u32 rx_page_size)
+static int
+fbnic_alloc_rx_qt_resources(struct fbnic_net *fbn,
+ struct fbnic_napi_vector *nv,
+ struct fbnic_q_triad *qt,
+ const struct netdev_queue_config *qcfg)
{
struct device *dev = fbn->netdev->dev.parent;
int err;
- err = fbnic_alloc_qt_page_pools(fbn, qt, qt->cmpl.q_idx, rx_page_size);
+ err = fbnic_alloc_qt_page_pools(fbn, qt, qt->cmpl.q_idx, qcfg);
if (err)
return err;
- fbnic_bdq_set_page_size(&qt->sub1, rx_page_size);
+ fbnic_bdq_set_page_size(&qt->sub1, qcfg->rx_page_size);
err = xdp_rxq_info_reg(&qt->xdp_rxq, fbn->netdev, qt->cmpl.q_idx,
nv->napi.napi_id);
@@ -2135,8 +2258,7 @@ static int fbnic_alloc_nv_resources(struct fbnic_net *fbn,
struct netdev_queue_config qcfg;
netdev_queue_config(fbn->netdev, nv->qt[i].cmpl.q_idx, &qcfg);
- err = fbnic_alloc_rx_qt_resources(fbn, nv, &nv->qt[i],
- qcfg.rx_page_size);
+ err = fbnic_alloc_rx_qt_resources(fbn, nv, &nv->qt[i], &qcfg);
if (err)
goto free_qt_resources;
}
@@ -2531,8 +2653,7 @@ static void fbnic_nv_fill(struct fbnic_napi_vector *nv)
struct fbnic_q_triad *qt = &nv->qt[t];
/* Populate the header and payload BDQs */
- fbnic_fill_bdq(&qt->sub0);
- fbnic_fill_bdq(&qt->sub1);
+ fbnic_fill_bdqs(qt);
}
}
@@ -2660,8 +2781,14 @@ static void fbnic_config_drop_mode_rcq(struct fbnic_napi_vector *nv,
struct fbnic_ring *rcq, bool tx_pause,
bool hdr_split)
{
+ struct fbnic_q_triad *qt = container_of(rcq, struct fbnic_q_triad,
+ cmpl);
struct fbnic_net *fbn = netdev_priv(nv->napi.dev);
- u32 drop_mode, rcq_ctl;
+ u32 drop_mode, hroom, rcq_ctl;
+
+ /* The maximum field value rounds up to one aligned unit more. */
+ hroom = min_t(u32, qt->rx_hroom,
+ FIELD_MAX(FBNIC_QUEUE_RDE_CTL0_MIN_HROOM_MASK));
if (!tx_pause && fbn->num_rx_queues > 1)
drop_mode = FBNIC_QUEUE_RDE_CTL0_DROP_IMMEDIATE;
@@ -2670,25 +2797,40 @@ static void fbnic_config_drop_mode_rcq(struct fbnic_napi_vector *nv,
/* Specify packet layout */
rcq_ctl = FIELD_PREP(FBNIC_QUEUE_RDE_CTL0_DROP_MODE_MASK, drop_mode) |
- FIELD_PREP(FBNIC_QUEUE_RDE_CTL0_MIN_HROOM_MASK, FBNIC_RX_HROOM) |
+ FIELD_PREP(FBNIC_QUEUE_RDE_CTL0_MIN_HROOM_MASK, hroom) |
FIELD_PREP(FBNIC_QUEUE_RDE_CTL0_MIN_TROOM_MASK, FBNIC_RX_TROOM) |
FIELD_PREP(FBNIC_QUEUE_RDE_CTL0_EN_HDR_SPLIT, hdr_split);
fbnic_ring_wr32(rcq, FBNIC_QUEUE_RDE_CTL0, rcq_ctl);
}
+static u32 fbnic_rcq_hds_thresh(const struct fbnic_net *fbn,
+ const struct fbnic_ring *rcq)
+{
+ const struct fbnic_q_triad *qt;
+ u32 hds_thresh = fbn->hds_thresh;
+
+ /* Link-up calls this without the netdev lock. Use the headroom
+ * cached in the queue; the pool may be going away.
+ */
+ qt = container_of(rcq, const struct fbnic_q_triad, cmpl);
+ if (qt->xsk_pool)
+ hds_thresh = fbnic_xsk_hds_thresh(qt->rx_hroom, hds_thresh);
+
+ return hds_thresh;
+}
+
void fbnic_config_drop_mode(struct fbnic_net *fbn, bool txp)
{
- bool hds;
int i, t;
- hds = fbn->hds_thresh < FBNIC_HDR_BYTES_MIN;
-
for (i = 0; i < fbn->num_napi; i++) {
struct fbnic_napi_vector *nv = fbn->napi[i];
for (t = 0; t < nv->rxt_count; t++) {
struct fbnic_q_triad *qt = &nv->qt[nv->txt_count + t];
+ u32 hds_thresh = fbnic_rcq_hds_thresh(fbn, &qt->cmpl);
+ bool hds = hds_thresh < FBNIC_HDR_BYTES_MIN;
fbnic_config_drop_mode_rcq(nv, &qt->cmpl, txp, hds);
}
@@ -2755,26 +2897,38 @@ void fbnic_config_rx_frames(struct fbnic_napi_vector *nv)
static void fbnic_enable_rcq(struct fbnic_napi_vector *nv,
struct fbnic_ring *rcq)
{
+ struct fbnic_q_triad *qt = container_of(rcq, struct fbnic_q_triad,
+ cmpl);
struct fbnic_net *fbn = netdev_priv(nv->napi.dev);
+ u32 payld_pg_cl = FBNIC_RX_PAYLD_PG_CL;
+ u32 payld_off = FBNIC_RX_PAYLD_OFFSET;
u32 log_size = fls(rcq->size_mask);
u32 rcq_ctl = 0;
bool hdr_split;
u32 hds_thresh;
+ if (!page_pool_supports_frag(qt->sub1.page_pool)) {
+ /* Align and retire each payload provider frame after one
+ * descriptor.
+ */
+ payld_pg_cl = ilog2(FBNIC_BD_PAGE_SIZE / FBNIC_RX_PAYLD_ALIGN);
+ payld_off = ilog2(FBNIC_BD_PAGE_SIZE / FBNIC_RX_PAYLD_ALIGN);
+ }
+
/* Force lower bound on MAX_HEADER_BYTES. Below this, all frames should
* be split at L4. It would also result in the frames being split at
* L2/L3 depending on the frame size.
*/
- hdr_split = fbn->hds_thresh < FBNIC_HDR_BYTES_MIN;
+ hds_thresh = fbnic_rcq_hds_thresh(fbn, rcq);
+ hdr_split = hds_thresh < FBNIC_HDR_BYTES_MIN;
fbnic_config_drop_mode_rcq(nv, rcq, fbn->tx_pause, hdr_split);
- hds_thresh = max(fbn->hds_thresh, FBNIC_HDR_BYTES_MIN);
+ hds_thresh = max(hds_thresh, FBNIC_HDR_BYTES_MIN);
rcq_ctl |= FIELD_PREP(FBNIC_QUEUE_RDE_CTL1_PADLEN_MASK, FBNIC_RX_PAD) |
FIELD_PREP(FBNIC_QUEUE_RDE_CTL1_MAX_HDR_MASK, hds_thresh) |
- FIELD_PREP(FBNIC_QUEUE_RDE_CTL1_PAYLD_OFF_MASK,
- FBNIC_RX_PAYLD_OFFSET) |
+ FIELD_PREP(FBNIC_QUEUE_RDE_CTL1_PAYLD_OFF_MASK, payld_off) |
FIELD_PREP(FBNIC_QUEUE_RDE_CTL1_PAYLD_PG_CL_MASK,
- FBNIC_RX_PAYLD_PG_CL);
+ payld_pg_cl);
fbnic_ring_wr32(rcq, FBNIC_QUEUE_RDE_CTL1, rcq_ctl);
/* Reset head/tail */
@@ -2931,8 +3085,7 @@ static int fbnic_queue_mem_alloc(struct net_device *dev,
struct fbnic_napi_vector *nv;
if (!netif_running(dev))
- return fbnic_alloc_qt_page_pools(fbn, qt, idx,
- qcfg->rx_page_size);
+ return fbnic_alloc_qt_page_pools(fbn, qt, idx, qcfg);
/* A failed PCIe recovery or resume can leave the datapath torn down
* while netif_running() is still true. This ndo runs before
@@ -2952,7 +3105,7 @@ static int fbnic_queue_mem_alloc(struct net_device *dev,
fbnic_ring_init(&qt->cmpl, real->cmpl.doorbell, real->cmpl.q_idx,
real->cmpl.flags);
- return fbnic_alloc_rx_qt_resources(fbn, nv, qt, qcfg->rx_page_size);
+ return fbnic_alloc_rx_qt_resources(fbn, nv, qt, qcfg);
}
static void fbnic_default_qcfg(struct net_device *dev,
@@ -2961,6 +3114,20 @@ static void fbnic_default_qcfg(struct net_device *dev,
qcfg->rx_page_size = PAGE_SIZE;
}
+/* Header starts are rounded up to the hardware alignment, and XDP needs its
+ * own headroom.
+ */
+static bool fbnic_rx_headroom_valid(u32 hroom)
+{
+ u32 max_hroom;
+
+ max_hroom = ALIGN(FIELD_MAX(FBNIC_QUEUE_RDE_CTL0_MIN_HROOM_MASK),
+ FBNIC_RX_HROOM_PAD);
+
+ return hroom >= XDP_PACKET_HEADROOM && hroom <= max_hroom &&
+ IS_ALIGNED(hroom, FBNIC_RX_HROOM_PAD);
+}
+
static int fbnic_validate_qcfg(struct net_device *dev,
struct netdev_queue_config *qcfg,
struct netlink_ext_ack *extack)
@@ -2989,6 +3156,11 @@ static int fbnic_validate_qcfg(struct net_device *dev,
return -EINVAL;
}
+ if (qcfg->rx_headroom && !fbnic_rx_headroom_valid(qcfg->rx_headroom)) {
+ NL_SET_ERR_MSG_MOD(extack, "rx_headroom is not supported");
+ return -EINVAL;
+ }
+
/* Payload fragments occupy multiples of FBNIC_RX_PAYLD_ALIGN bytes.
* Keep at least one reference in the bias until fbnic_clean_bdq()
* observes a completion from a subsequent allocation.
@@ -3119,5 +3291,5 @@ const struct netdev_queue_mgmt_ops fbnic_queue_mgmt_ops = {
.ndo_queue_stop = fbnic_queue_stop,
.ndo_default_qcfg = fbnic_default_qcfg,
.ndo_validate_qcfg = fbnic_validate_qcfg,
- .supported_params = QCFG_RX_PAGE_SIZE,
+ .supported_params = QCFG_RX_PAGE_SIZE | QCFG_RX_HEADROOM,
};
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h
index 6ee3cce3d942..abc96a12ea5c 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h
@@ -15,6 +15,7 @@
#include "fbnic_csr.h"
struct fbnic_net;
+struct xsk_buff_pool;
/* Guarantee we have space needed for storing the buffer
* To store the buffer we need:
@@ -77,12 +78,14 @@ struct fbnic_net;
(4096 - FBNIC_RX_HROOM - FBNIC_RX_TROOM - FBNIC_RX_PAD)
#define FBNIC_HDS_THRESH_DEFAULT \
(1536 - FBNIC_RX_PAD)
+#define FBNIC_XSK_HDS_THRESH 3072
#define FBNIC_HDR_BYTES_MIN 256
struct fbnic_pkt_buff {
struct xdp_buff buff;
ktime_t hwtstamp;
bool add_frag_failed;
+ struct skb_shared_info sinfo;
};
struct fbnic_queue_stats {
@@ -162,6 +165,9 @@ struct fbnic_ring {
struct fbnic_q_triad {
struct fbnic_ring sub0, sub1, cmpl;
struct xdp_rxq_info xdp_rxq;
+ struct xsk_buff_pool *xsk_pool;
+ u16 rx_hroom;
+ u16 rx_frame_off;
};
struct fbnic_napi_vector {
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 14/15] eth: fbnic: Support AF_XDP zero-copy transmit
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (12 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 13/15] eth: fbnic: Support AF_XDP zero-copy receive Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
2026-10-02 19:00 ` [RFC net-next 15/15] Documentation: xsk: Document page-pool zero copy Björn Töpel
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
AF_XDP zero copy cannot be advertised until fbnic can send socket TX
descriptors and report their completion.
Send socket descriptors from the RX NAPI poll, and add
ndo_xsk_wakeup to schedule that poll. Record the type of each XDP TX
ring entry. On completion, page-pool buffers go back to page_pool
and UMEM frames go to the AF_XDP completion ring. XDP_TX of provider
RX buffers uses the same records.
Each poll sends at most 64 descriptors per socket, so the
application and NAPI can run at the same time. When a vector has
several sockets, the first one served changes on each poll. The
batch helper does not take a multi-buffer packet that does not fit
the budget left. If a batch was sent and the budget left is smaller
than a maximum-size packet, keep NAPI scheduled.
Socket TX shares the XDP TX ring with XDP_TX. Do socket TX after RX:
XDP_TX drops a frame that does not fit, but a socket descriptor
waits in the socket TX ring.
Advertise AF_XDP zero copy on systems with 4 KiB pages, now that RX
and TX both work.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
.../net/ethernet/meta/fbnic/fbnic_netdev.c | 30 +++
drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 209 ++++++++++++++++--
drivers/net/ethernet/meta/fbnic/fbnic_txrx.h | 18 ++
3 files changed, 238 insertions(+), 19 deletions(-)
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c b/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c
index 038e91ef14e4..f8495b665d97 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_netdev.c
@@ -666,6 +666,31 @@ static int fbnic_bpf(struct net_device *netdev, struct netdev_bpf *bpf)
}
}
+static int fbnic_xsk_wakeup(struct net_device *netdev, u32 queue_id,
+ u32 __always_unused flags)
+{
+ struct fbnic_net *fbn = netdev_priv(netdev);
+ struct xsk_buff_pool *pool;
+ struct napi_struct *napi;
+
+ if (queue_id >= fbn->num_rx_queues)
+ return -EINVAL;
+ if (!netif_running(netdev))
+ return -ENETDOWN;
+ pool = xsk_get_pool_from_rxq(netdev, queue_id);
+ if (!pool)
+ return -EINVAL;
+
+ napi = __netif_get_rx_queue(netdev, queue_id)->napi;
+ if (WARN_ON_ONCE(!napi))
+ return -EINVAL;
+
+ if (!napi_if_scheduled_mark_missed(napi))
+ napi_schedule(napi);
+
+ return 0;
+}
+
static const struct net_device_ops fbnic_netdev_ops = {
.ndo_open = fbnic_open,
.ndo_stop = fbnic_stop,
@@ -677,6 +702,7 @@ static const struct net_device_ops fbnic_netdev_ops = {
.ndo_set_rx_mode_async = fbnic_set_rx_mode,
.ndo_get_stats64 = fbnic_get_stats64,
.ndo_bpf = fbnic_bpf,
+ .ndo_xsk_wakeup = fbnic_xsk_wakeup,
.ndo_hwtstamp_get = fbnic_hwtstamp_get,
.ndo_hwtstamp_set = fbnic_hwtstamp_set,
};
@@ -935,6 +961,10 @@ struct net_device *fbnic_netdev_alloc(struct fbnic_dev *fbd)
netdev->xdp_features = NETDEV_XDP_ACT_BASIC |
NETDEV_XDP_ACT_REDIRECT |
NETDEV_XDP_ACT_RX_SG;
+ if (PAGE_SIZE == FBNIC_BD_PAGE_SIZE) {
+ netdev->xdp_features |= NETDEV_XDP_ACT_XSK_ZEROCOPY;
+ netdev->xdp_zc_max_segs = FBNIC_MAX_RX_DATA_DESC;
+ }
netdev->min_mtu = IPV6_MIN_MTU;
netdev->max_mtu = FBNIC_MAX_JUMBO_FRAME_SIZE - ETH_HLEN;
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
index e6043607d4e2..6990a7e0fe2f 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c
@@ -42,6 +42,7 @@ struct fbnic_xmit_cb {
#define FBNIC_XMIT_CB(__skb) ((struct fbnic_xmit_cb *)((__skb)->cb))
#define FBNIC_XMIT_NOUNMAP ((void *)1)
+#define FBNIC_XSK_TX_BUDGET 64
u32 __iomem *fbnic_ring_csr_base(const struct fbnic_ring *ring)
{
@@ -657,22 +658,79 @@ static void fbnic_clean_twq0(struct fbnic_napi_vector *nv, int napi_budget,
}
}
-static void fbnic_clean_twq1(struct fbnic_napi_vector *nv, bool pp_allow_direct,
+struct fbnic_xsk_tx_complete {
+ struct xsk_buff_pool *pool;
+ u32 descs;
+};
+
+enum fbnic_pp_recycle_mode {
+ FBNIC_PP_RECYCLE_DEFERRED,
+ FBNIC_PP_RECYCLE_DIRECT,
+};
+
+static void
+fbnic_xsk_tx_complete_flush(struct fbnic_xsk_tx_complete *complete)
+{
+ if (!complete->descs)
+ return;
+
+ xsk_tx_completed(complete->pool, complete->descs);
+ complete->descs = 0;
+}
+
+static void
+fbnic_xsk_tx_complete_count(struct fbnic_xsk_tx_complete *complete,
+ struct xsk_buff_pool *pool)
+{
+ if (complete->pool != pool) {
+ fbnic_xsk_tx_complete_flush(complete);
+ complete->pool = pool;
+ }
+
+ complete->descs++;
+}
+
+static void
+fbnic_clean_twq1_buf(struct fbnic_xdp_tx_buf *buf,
+ struct fbnic_xsk_tx_complete *complete,
+ enum fbnic_pp_recycle_mode mode)
+{
+ switch (buf->type) {
+ case FBNIC_XDP_TX_BUF_NETMEM:
+ page_pool_put_full_netmem(netmem_get_pp(buf->netmem),
+ buf->netmem,
+ mode == FBNIC_PP_RECYCLE_DIRECT);
+ break;
+ case FBNIC_XDP_TX_BUF_XSK:
+ fbnic_xsk_tx_complete_count(complete, buf->xsk_pool);
+ break;
+ default:
+ WARN_ON_ONCE(buf->type != FBNIC_XDP_TX_BUF_NONE);
+ break;
+ }
+
+ buf->netmem = 0;
+ buf->type = FBNIC_XDP_TX_BUF_NONE;
+}
+
+static void fbnic_clean_twq1(struct fbnic_napi_vector *nv,
+ enum fbnic_pp_recycle_mode mode,
struct fbnic_ring *ring, bool discard,
unsigned int hw_head)
{
+ struct fbnic_xsk_tx_complete complete = { };
u64 total_bytes = 0, total_packets = 0;
unsigned int head = ring->head;
while (hw_head != head) {
- struct page *page;
+ struct fbnic_xdp_tx_buf *buf;
u64 twd;
if (unlikely(!(ring->desc[head] & FBNIC_TWD_TYPE(AL))))
goto next_desc;
twd = le64_to_cpu(ring->desc[head]);
- page = ring->tx_buf[head];
+ buf = &ring->xdp_buf[head];
/* TYPE_AL is 2, TYPE_LAST_AL is 3. So this trick gives
* us one increment per packet, with no branches.
@@ -681,13 +739,14 @@ static void fbnic_clean_twq1(struct fbnic_napi_vector *nv, bool pp_allow_direct,
FBNIC_TWD_TYPE_AL;
total_bytes += FIELD_GET(FBNIC_TWD_LEN_MASK, twd);
- page_pool_put_page(pp_page_to_nmdesc(page)->pp, page, -1,
- pp_allow_direct);
+ fbnic_clean_twq1_buf(buf, &complete, mode);
next_desc:
head++;
head &= ring->size_mask;
}
+ fbnic_xsk_tx_complete_flush(&complete);
+
if (!total_bytes)
return;
@@ -821,7 +880,8 @@ static void fbnic_clean_twq(struct fbnic_napi_vector *nv, int napi_budget,
if (head1 >= 0) {
qt->cmpl.deferred_head = -1;
if (napi_budget)
- fbnic_clean_twq1(nv, true, &qt->sub1, false, head1);
+ fbnic_clean_twq1(nv, FBNIC_PP_RECYCLE_DIRECT, &qt->sub1,
+ false, head1);
else
qt->cmpl.deferred_head = head1;
}
@@ -1203,7 +1263,7 @@ static long fbnic_pkt_tx(struct fbnic_napi_vector *nv,
unsigned int tail = ring->tail;
struct skb_shared_info *shinfo;
skb_frag_t *frag = NULL;
- struct page *page;
+ netmem_ref netmem;
dma_addr_t dma;
__le64 *twd;
@@ -1221,18 +1281,21 @@ static long fbnic_pkt_tx(struct fbnic_napi_vector *nv,
return -FBNIC_XDP_CONSUME;
}
- page = virt_to_page(pkt->buff.data_hard_start);
- offset = offset_in_page(pkt->buff.data);
- dma = page_pool_get_dma_addr(page);
-
+ netmem = xdp_buff_get_netmem(&pkt->buff);
+ offset = (u8 *)pkt->buff.data -
+ (u8 *)netmem_address(netmem);
+ dma = page_pool_get_dma_addr_netmem(netmem);
size = pkt->buff.data_end - pkt->buff.data;
while (nsegs--) {
+ struct fbnic_xdp_tx_buf *buf = &ring->xdp_buf[tail];
+
dma_sync_single_range_for_device(nv->dev, dma, offset, size,
DMA_BIDIRECTIONAL);
dma += offset;
- ring->tx_buf[tail] = page;
+ buf->netmem = netmem;
+ buf->type = FBNIC_XDP_TX_BUF_NETMEM;
twd = &ring->desc[tail];
*twd = cpu_to_le64(FIELD_PREP(FBNIC_TWD_ADDR_MASK, dma) |
@@ -1247,8 +1310,8 @@ static long fbnic_pkt_tx(struct fbnic_napi_vector *nv,
break;
offset = skb_frag_off(frag);
- page = skb_frag_page(frag);
- dma = page_pool_get_dma_addr(page);
+ netmem = skb_frag_netmem(frag);
+ dma = page_pool_get_dma_addr_netmem(netmem);
size = skb_frag_size(frag);
data_len -= size;
@@ -1490,6 +1553,102 @@ static int fbnic_clean_rcq(struct fbnic_napi_vector *nv,
return complete ? packets : budget;
}
+static bool fbnic_xsk_tx_pool(struct fbnic_ring *ring,
+ struct xsk_buff_pool *pool)
+{
+ unsigned int budget, tail = ring->tail;
+ struct xdp_desc *descs;
+ bool complete = true;
+ bool need_wakeup;
+ u32 i, count;
+
+ need_wakeup = xsk_uses_need_wakeup(pool);
+ if (need_wakeup && ring->head == ring->tail)
+ xsk_set_tx_need_wakeup(pool);
+
+ budget = min_t(unsigned int, fbnic_desc_unused(ring),
+ FBNIC_XSK_TX_BUDGET);
+ if (!budget)
+ goto wake;
+
+ count = xsk_tx_peek_release_desc_batch(pool, budget);
+ /* A shared UMEM bind can allocate the array after pool creation. */
+ descs = pool->tx_descs;
+ for (i = 0; i < count; i++) {
+ struct fbnic_xdp_tx_buf *buf = &ring->xdp_buf[tail];
+ struct xdp_desc_ctx ctx;
+ struct xdp_desc *desc;
+ unsigned int type;
+
+ desc = &descs[i];
+ ctx = xsk_buff_raw_get_ctx(pool, desc->addr, desc->options);
+ xsk_buff_raw_dma_sync_for_device(pool, ctx.dma, desc->len);
+
+ buf->xsk_pool = pool;
+ buf->type = FBNIC_XDP_TX_BUF_XSK;
+ type = xsk_is_eop_desc(desc) ? FBNIC_TWD_TYPE_LAST_AL :
+ FBNIC_TWD_TYPE_AL;
+ ring->desc[tail] =
+ cpu_to_le64(FIELD_PREP(FBNIC_TWD_ADDR_MASK, ctx.dma) |
+ FIELD_PREP(FBNIC_TWD_LEN_MASK, desc->len) |
+ FIELD_PREP(FBNIC_TWD_TYPE_MASK, type));
+
+ tail++;
+ tail &= ring->size_mask;
+ }
+
+ if (count)
+ ring->tail = tail;
+ /* The batch helper stops before a partial multi-buffer packet. Keep
+ * polling if the unused budget was too small for one maximum packet.
+ */
+ complete = !count || budget - count >= pool->xdp_zc_max_segs;
+
+wake:
+ if (need_wakeup && ring->head != ring->tail)
+ xsk_clear_tx_need_wakeup(pool);
+
+ return complete;
+}
+
+static bool fbnic_xsk_tx(struct fbnic_napi_vector *nv)
+{
+ struct fbnic_ring *ring = &nv->qt[0].sub1;
+ unsigned int tail = ring->tail;
+ unsigned int i, queue;
+ bool complete = true;
+
+ if (!nv->rxt_count)
+ return true;
+
+ queue = nv->xsk_tx_start;
+ do {
+ struct xsk_buff_pool *pool;
+
+ i = nv->txt_count + queue;
+ pool = nv->qt[i].xsk_pool;
+ if (pool)
+ complete &= fbnic_xsk_tx_pool(ring, pool);
+
+ queue++;
+ if (queue == nv->rxt_count)
+ queue = 0;
+ } while (queue != nv->xsk_tx_start);
+
+ nv->xsk_tx_start++;
+ if (nv->xsk_tx_start == nv->rxt_count)
+ nv->xsk_tx_start = 0;
+
+ if (tail != ring->tail) {
+ /* Force DMA writes to flush before writing to tail */
+ dma_wmb();
+
+ writel(ring->tail, ring->doorbell);
+ }
+
+ return complete;
+}
+
static void fbnic_nv_irq_disable(struct fbnic_napi_vector *nv)
{
struct fbnic_dev *fbd = nv->fbd;
@@ -1512,6 +1671,7 @@ static int fbnic_poll(struct napi_struct *napi, int budget)
struct fbnic_napi_vector *nv = container_of(napi,
struct fbnic_napi_vector,
napi);
+ bool xsk_complete = true;
int i, j, work_done = 0;
for (i = 0; i < nv->txt_count; i++)
@@ -1520,7 +1680,14 @@ static int fbnic_poll(struct napi_struct *napi, int budget)
for (j = 0; j < nv->rxt_count; j++, i++)
work_done += fbnic_clean_rcq(nv, &nv->qt[i], budget);
- if (work_done >= budget)
+ /* XDP_TX shares this ring and drops frames that do not fit, while
+ * socket TX descriptors wait in the socket's TX ring. Serve XDP_TX
+ * first.
+ */
+ if (budget && nv->rxt_count)
+ xsk_complete = fbnic_xsk_tx(nv);
+
+ if (work_done >= budget || !xsk_complete)
return budget;
if (likely(napi_complete_done(napi, work_done)))
@@ -1884,7 +2051,8 @@ static int fbnic_alloc_napi_vector(struct fbnic_dev *fbd, struct fbnic_net *fbn,
if (xdp_count > 0) {
unsigned int xdp_idx = FBNIC_MAX_TXQS + rxq_idx;
- fbnic_ring_init(&qt->sub1, db, xdp_idx, flags);
+ fbnic_ring_init(&qt->sub1, db, xdp_idx,
+ flags | FBNIC_RING_F_XDP);
fbn->tx[xdp_idx] = &qt->sub1;
xdp_count--;
} else {
@@ -2031,9 +2199,12 @@ static int fbnic_alloc_tx_ring_buffer(struct fbnic_ring *txr)
{
size_t size = array_size(sizeof(*txr->tx_buf), txr->size_mask + 1);
- txr->tx_buf = kvzalloc(size, GFP_KERNEL | __GFP_NOWARN);
+ if (txr->flags & FBNIC_RING_F_XDP)
+ size = array_size(sizeof(*txr->xdp_buf), txr->size_mask + 1);
- return txr->tx_buf ? 0 : -ENOMEM;
+ txr->buffer = kvzalloc(size, GFP_KERNEL | __GFP_NOWARN);
+
+ return txr->buffer ? 0 : -ENOMEM;
}
static int fbnic_alloc_tx_ring_resources(struct fbnic_net *fbn,
@@ -2602,7 +2773,7 @@ static void fbnic_nv_flush(struct fbnic_napi_vector *nv)
/* Clean the work queues of unprocessed work */
fbnic_clean_twq0(nv, 0, &qt->sub0, true, qt->sub0.tail);
- fbnic_clean_twq1(nv, false, &qt->sub1, true,
+ fbnic_clean_twq1(nv, FBNIC_PP_RECYCLE_DEFERRED, &qt->sub1, true,
qt->sub1.tail);
/* Reset completion queue descriptor ring */
diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h
index abc96a12ea5c..2fcfe6ac0232 100644
--- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h
+++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.h
@@ -36,6 +36,7 @@ struct xsk_buff_pool;
* + 4 descriptors for payload
*/
#define FBNIC_MAX_RX_PKT_DESC 7
+#define FBNIC_MAX_RX_DATA_DESC (FBNIC_MAX_RX_PKT_DESC - 2)
#define FBNIC_RX_DESC_MIN roundup_pow_of_two(FBNIC_MAX_RX_PKT_DESC * 2)
#define FBNIC_MAX_TXQS 128u
@@ -73,6 +74,7 @@ struct xsk_buff_pool;
#define FBNIC_RING_F_DISABLED BIT(0)
#define FBNIC_RING_F_CTX BIT(1)
#define FBNIC_RING_F_STATS BIT(2) /* Ring's stats may be used */
+#define FBNIC_RING_F_XDP BIT(3)
#define FBNIC_HDS_THRESH_MAX \
(4096 - FBNIC_RX_HROOM - FBNIC_RX_TROOM - FBNIC_RX_PAD)
@@ -121,11 +123,26 @@ struct fbnic_rx_buf {
long pagecnt_bias;
};
+enum fbnic_xdp_tx_buf_type {
+ FBNIC_XDP_TX_BUF_NONE = 0,
+ FBNIC_XDP_TX_BUF_NETMEM,
+ FBNIC_XDP_TX_BUF_XSK,
+};
+
+struct fbnic_xdp_tx_buf {
+ union {
+ struct xsk_buff_pool *xsk_pool;
+ netmem_ref netmem;
+ };
+ u8 type;
+};
+
struct fbnic_ring {
/* Pointer to buffer specific info */
union {
struct fbnic_pkt_buff *pkt; /* RCQ */
struct fbnic_rx_buf *rx_buf; /* BDQ */
+ struct fbnic_xdp_tx_buf *xdp_buf; /* TWQ1 */
void **tx_buf; /* TWQ */
void *buffer; /* Generic pointer */
};
@@ -179,6 +196,7 @@ struct fbnic_napi_vector {
u16 v_idx;
u8 txt_count;
u8 rxt_count;
+ u8 xsk_tx_start;
struct fbnic_q_triad qt[];
};
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [RFC net-next 15/15] Documentation: xsk: Document page-pool zero copy
2026-10-02 19:00 [RFC net-next 00/15] xsk: Zero copy through page-pool memory providers Björn Töpel
` (13 preceding siblings ...)
2026-10-02 19:00 ` [RFC net-next 14/15] eth: fbnic: Support AF_XDP zero-copy transmit Björn Töpel
@ 2026-10-02 19:00 ` Björn Töpel
14 siblings, 0 replies; 17+ messages in thread
From: Björn Töpel @ 2026-10-02 19:00 UTC (permalink / raw)
To: Magnus Karlsson, Maciej Fijalkowski, Stanislav Fomichev,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Alexander Duyck, kernel-team, Andrew Lunn, Jesper Dangaard Brouer,
Ilias Apalodimas, Alexei Starovoitov, Daniel Borkmann,
John Fastabend, Pavel Begunkov, Jens Axboe, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis,
Ihor Solodrai, netdev, bpf, io-uring
Cc: Björn Töpel, Mike Marciniszyn (Meta), Weiming Shi,
Nikolay Aleksandrov, David Wei, Alexander Lobakin, linux-doc,
linux-kernel, Mina Almasry
Describe AF_XDP zero copy on page-pool drivers: the aligned 4 KiB
UMEM requirement, how the headroom reaches the driver through the
queue configuration, the device-chosen offset in fragment
descriptors, the scatter-gather layout, buffer ownership, provider
locking, and the order in which a driver enables it.
Signed-off-by: Björn Töpel <bjorn@kernel.org>
---
Documentation/networking/af_xdp.rst | 65 +++++++++++++++++++++++++++++
1 file changed, 65 insertions(+)
diff --git a/Documentation/networking/af_xdp.rst b/Documentation/networking/af_xdp.rst
index cc3f0d16b28f..bdf4029aedeb 100644
--- a/Documentation/networking/af_xdp.rst
+++ b/Documentation/networking/af_xdp.rst
@@ -346,6 +346,71 @@ Note that a UMEM can be shared between sockets on the same queue id
and device, as well as between queues on the same device and between
devices at the same time.
+Page-pool backed zero-copy
+--------------------------
+
+Drivers which use the queue management API can obtain UMEM frames through a
+page-pool memory provider. This is selected by the driver when zero-copy mode
+is requested and does not require a new userspace flag. The FILL ring remains
+the source of receive buffers and all normal AF_XDP ownership rules apply.
+Once installed, the driver remains an ordinary page-pool consumer: allocation,
+DMA synchronization, recycling, release, and refill use the normal page-pool
+interfaces, while provider callbacks hide the UMEM-specific operations.
+
+The initial provider requires 4 KiB base pages and aligned 4 KiB chunks.
+Unaligned chunks are rejected. Configured UMEM headroom is supported. The
+provider requests it through the queue configuration, and the driver includes
+it in the receive DMA offset; with multi-buffer packets it applies to the first
+descriptor as described below.
+
+Like the normal XSK buffer allocator, provider-backed page-pool allocation and
+recycling run in the receive queue's NAPI context. A queue restart prepares
+its replacement before it stops the current queue, so two page pools can use
+one provider for a short time. A provider lock serializes FILL-ring
+consumption and the provider's buffer stack. It is taken once per page-pool
+refill of up to 64 buffers, not per packet. Generic XDP cannot redirect to a
+provider-backed socket; its copy-mode receive path retains the existing XSK
+receive lock.
+
+UMEM frames retain the direct XSK ownership model. The provider does not add a
+per-frame reference count, generation, ownership bitmap, or quarantine state.
+A frame moves between the FILL ring, the owning NAPI context, and userspace;
+userspace must not publish a frame which it does not own. Page-pool teardown
+accounting protects the lifetime of the pool, not ownership of an individual
+UMEM frame.
+
+Provider-backed buffers do not leave that context as kernel-owned memory.
+``XDP_PASS`` copies the packet to kernel-backed skb storage before returning
+the UMEM buffers. Redirects other than a compatible XSKMAP target likewise
+copy to kernel memory. A compatible XSKMAP transfer publishes the UMEM
+descriptors directly to userspace; returning them through the FILL ring makes
+them available to the same queue context again.
+
+For multi-buffer packets, fragment descriptors are assembled in transient
+kernel-owned storage belonging to the RX queue. They are never stored in the
+user-writable UMEM, and are consumed before the NAPI context starts the next
+packet.
+
+The copy on ``XDP_PASS`` is intentional: an skb may outlive the receive NAPI
+poll, whereas a provider frame must be returned by the context which allocated
+it. Applications which expect most packets to pass to the network stack should
+therefore account for this copy when choosing page-pool backed zero-copy.
+
+Drivers may impose additional layout and queue requirements. The initial fbnic
+support accepts configured UMEM headroom from 0 through 256 bytes in 128-byte
+increments (up to 512 bytes including ``XDP_PACKET_HEADROOM``).
+Packets larger than its selected header-data-split threshold require an
+``XDP_USE_SG`` socket and an XDP program with fragment support. Their
+continuation descriptors start at offsets chosen by the device.
+When these restrictions are not met, a bind forced with ``XDP_ZEROCOPY``
+fails with an error; automatic mode may fall back to copy mode.
+
+On a running device, installing or removing the provider restarts the
+selected hardware queue.
+Applications should populate the FILL ring before binding when possible. If
+the ring is empty, the kernel schedules the queue once after installation;
+the usual ``XDP_USE_NEED_WAKEUP`` rules apply after that.
+
XDP_USE_NEED_WAKEUP bind flag
-----------------------------
--
2.55.0
^ permalink raw reply related [flat|nested] 17+ messages in thread