From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj2-f8.google.com (mail-pj2-f8.google.com [74.125.227.136]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 333B54C33D8 for ; Mon, 5 Oct 2026 16:51:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.136 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791219082; cv=none; b=ckFZ2b4uXzCJ22mGbt1UmefnYbNbjbNSgIdBwagr+gDpAg7moa3eO+2lI8AkWKD4KY+a8dqRHu3PgbttD0ovS/NQyPDQYJ8EOfRc1ScFfvBoqKJgQICiMJQzgitP1XQkc5OUrbSAXGpSLBT8dPy/2l//boDdcmIgmGLfRzPHNKs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791219082; c=relaxed/simple; bh=ZtOtfI8xTZljtzEITNuddg7KMK3dN7NLgg7QIotbHg0=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=oQvUvsoQTcvJSJa41N7ECWbyw8DYjKvsRiV7vFVBBqmtpW+B2HdTZYdkO3nmp12mW6Ufz4viSeNqTGZj0Ku2iClD8HYEEIY2yRjTNqozB04tiCmo15C4HWkeccsfE7qV40As7nefK5vn8ZwtJcVDimEFgAcJDqLbiuU9hsiZto4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jYBEAguj; arc=none smtp.client-ip=74.125.227.136 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jYBEAguj" Received: by mail-pj2-f8.google.com with SMTP id 98e67ed59e1d1-398b9f722abso1144288a91.0 for ; Mon, 05 Oct 2026 09:51:21 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1791219080; x=1791823880; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=b4mkDJheIxHRdMD7TqVn4ykGrzx67Q/w28iQIL9i/AE=; b=jYBEAguj/HXC2O9ml4GkNN/+9pMpbTf40fb8JKLkjR7FD8OKSyr70p61oJulOnjdra kQaMEQjqyv7sgrI2g2fg6PPmcYTR8lSmBlH090Iui9CcaC+IW37lvow5Ppb4NrXhAkO2 IUUbQTvR8oSRFHXnBVOn4p7dae5R/an/Hk/XbBVSAfbNwhQWqhTg9eyVfwJuFg/U1L7j TW4UrRlz0bYBlxFVVvbKXXheqRt+HGKxQJJjuqIluT7hvwbwkHHQEyKPKJCKrRlbR9jh fqoadKhkTSprHmHkHPhqDaUgY3EKNIo4K78uf9PO/TOVsW6IOOw9qLPh+UHkye4inhDv KQvw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791219080; x=1791823880; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=b4mkDJheIxHRdMD7TqVn4ykGrzx67Q/w28iQIL9i/AE=; b=gAHcOai8X9AhxyHQyUJRUv/Lzhf3hUodicmHC+yDjln/GqOtFd1GkLCAeGHoflHjoY yrrq1IeeiXO1zUJtgKXvEbBeVC3ftGHtU6dLN3Go/vvN5CNo6IbVVPQcwRnZYn3pp0UA IL88wYQhE11tCFtzoGGEkw3N6Yqlm3ckZFwofIrGmucUII1OQ1nT3ZI7yTA/PFa9Rttk g5N3ZTeUpHwO8baNQeoeSEDxL9CZUt63NncdFiogzLTUEhv9px38e6zkEtZhcY0lqDHK 5OSE0ezZagafzqGvru+W7hSuK0G5FNCXvLSPPLVHNH7Wa++2d7zqHK/6Ia1kYfoCONja A+7g== X-Forwarded-Encrypted: i=1; AKwUvBxCpeZZnvLpf1YZX6TyxSHM9odRizxwfiPfTmK/8CRuqhInRLdp33sWrVldATr5M94vncSgJ8akDQ==@vger.kernel.org X-Gm-Message-State: AFq9FYISeM4XP+/fABhhtWtEurzekx3cbrxIGKrRZIF7korupna/VBbP 67iD8xhQ9REWr/I4dw+2i6/xl5+8zn9dqy9Nt7a9sHCPS5mArUFdi5V4 X-Gm-Gg: AYBFou3izVGtY7+FsZ2o/YoDctFdDA4H4gPjElfS/r2sFfc2NUEwiGD4QJJT2FeHD/q 3f7rPnRTdChJQh7+nk1EUp/cBy3VKHlPv9uxy6eZzhD++hab3hTvC9ZR8XZrZ1qTyyVOlD1swZd GKOXP2SeocP0yvzOoHgl9rcvWMGbQAbdj2v33h6jpeLafvlPoD3Q9byFv/VhbKu3WC+tjGmFeS6 yTlclYnv95VUq5oDSlA6aSIuT8AEPBZHcZx1mD9gY+459ZpUFLyfo0cwPX+k07Tm80dc/v/sjrw zyOIdqWXmE2G3VHRuPk99wQTc0Yq1eGNd/IfSVc1c01dj9Xa1+oXbEzi0twMjO489/dmqnKgiBG ZLDAkVDPAKNo/EMQ5ilLvPoNRam9WrdR0gCUeCfjw+iVpVpXV3pt/6lIIIAytadR0o7RYQU4Ec7 DUR7UX1DAHY4Kanttkxqf/9OtYQIFw8DuLqWYZE6F6VR29hLh4vFNZefPIDru9PNw= X-Received: by 2002:a17:90b:2ec3:b0:3a6:fec5:c184 with SMTP id 98e67ed59e1d1-3a6fec5d190mr6627900a91.7.1791219080418; Mon, 05 Oct 2026 09:51:20 -0700 (PDT) Received: from localhost ([2a03:2880:2ff:1::]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3a851525bffsm261628a91.0.2026.10.05.09.51.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 05 Oct 2026 09:51:20 -0700 (PDT) Date: Mon, 5 Oct 2026 09:51:15 -0700 From: Stanislav Fomichev To: =?utf-8?B?QmrDtnJuIFTDtnBlbA==?= Cc: Magnus Karlsson , Maciej Fijalkowski , Stanislav Fomichev , "David S. Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Jonathan Corbet , Shuah Khan , Randy Dunlap , Alexander Duyck , kernel-team@meta.com, Andrew Lunn , Jesper Dangaard Brouer , Ilias Apalodimas , Alexei Starovoitov , Daniel Borkmann , John Fastabend , Pavel Begunkov , Jens Axboe , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , netdev@vger.kernel.org, bpf@vger.kernel.org, io-uring@vger.kernel.org, "Mike Marciniszyn (Meta)" , Weiming Shi , Nikolay Aleksandrov , David Wei , Alexander Lobakin , linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, Mina Almasry Subject: Re: [RFC net-next 05/15] page_pool: Extend memory provider operations Message-ID: References: <20261002190018.696925-1-bjorn@kernel.org> <20261002190018.696925-6-bjorn@kernel.org> Precedence: bulk X-Mailing-List: io-uring@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <20261002190018.696925-6-bjorn@kernel.org> On 10/02, Björn Töpel wrote: > The current memory providers give memory that the CPU cannot read, > and they can always refill. An AF_XDP provider is different. The CPU > can read its memory, it has only the buffers that userspace puts in > the FILL ring, and its buffers go back to userspace instead of being > freed. It needs four things from page_pool: > > - A CPU address for each net_iov. A provider can set the address of > the first net_iov in its net_iov_area and the size each net_iov > covers. netmem_address() then computes the address without a call > into the provider; drivers call it for every packet. A > net_iov_area without an address stays unreadable. > page_pool_is_unreadable() is false for a readable provider. > > - A way to say that the refill is not done. The new refill_done > callback tells the driver whether it may stop refilling. A > provider that still has buffers returns false, and the driver > keeps NAPI scheduled. > > - A way to give up a buffer without freeing it. Add a batched call > that ends page_pool ownership of buffers that leave through the > provider, for example to userspace. It clears the page_pool link, > so a later page pool can take the buffer. Add the matching batched > call that sets the link, and a DMA sync helper for drivers that > post a buffer again directly. > > - No buffer splitting, because a split buffer has several owners. > Add MP_CAP_FRAG. page_pool refuses fragment allocation from a > provider without it. > > devmem and io_uring set MP_CAP_FRAG and keep their fragment > behaviour. Their refill paths now use the batched call that sets the > link. > > Signed-off-by: Björn Töpel > --- > include/net/netmem.h | 31 ++++++++++- > include/net/page_pool/helpers.h | 62 +++++++++++++++++++++- > include/net/page_pool/memory_provider.h | 28 +++++----- > io_uring/zcrx.c | 4 +- > net/core/devmem.c | 7 ++- > net/core/page_pool.c | 68 ++++++++++++++++++++----- > 6 files changed, 165 insertions(+), 35 deletions(-) > > diff --git a/include/net/netmem.h b/include/net/netmem.h > index cc97611632dc..e6dff0b01581 100644 > --- a/include/net/netmem.h > +++ b/include/net/netmem.h > @@ -105,6 +105,12 @@ struct net_iov_area { > [..] > /* Offset into the dma-buf where this chunk starts. */ > unsigned long base_virtual; Will send a patch to drop this one, don't think we use it anymore.. [..] > + /* CPU address of the first net_iov's memory, or NULL when the CPU > + * cannot read the area. Each net_iov covers 1 << @niov_shift bytes. > + */ > + void *vaddr; Who is setting this? Somewhere in another patch? Can we keep it here? > + u8 niov_shift; If we need niov_shift at area level, let's move the one from binding here? Then the existing callers can do binding->area.niov_shift. > }; > > static inline struct net_iov_area *net_iov_owner(const struct net_iov *niov) > @@ -117,6 +123,23 @@ static inline unsigned int net_iov_idx(const struct net_iov *niov) > return niov - net_iov_owner(niov)->niovs; > } > > +static inline bool net_iov_is_readable(const struct net_iov *niov) > +{ > + return net_iov_owner(niov)->vaddr; > +} > + > +static inline void *net_iov_address(const struct net_iov *niov) > +{ > + const struct net_iov_area *area = net_iov_owner(niov); > + unsigned long off; > + > + if (!area->vaddr) > + return NULL; > + > + off = (unsigned long)net_iov_idx(niov) << area->niov_shift; > + return area->vaddr + off; Or, alternatively, store dma_base in area and do: return niov->dma_addr - area.dma_base + vadd ?