From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f47.google.com (mail-wm1-f47.google.com [209.85.128.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CDDCD347533 for ; Wed, 7 Oct 2026 01:32:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.47 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791336740; cv=none; b=fHmiZpZvtOVPPDi8EjDGbDOmProtCWTGWIhZTo2jczu34nWDSX82RIV+q0FbHGCW1ZoT0XRBPB6DQCqYhk6A9HpR4S5D1bMhLLhrZopbqJ0GzAOE1Lv8xxoNUlWj4O1mInZmxYHL36IfAaLwMS6aGDmb9nblyRyXeGF/4gyxi/U= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791336740; c=relaxed/simple; bh=zJgertkyenaDyaojdQoZusF5JfnPhGy9EVpiIwz/cT8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aWaXe73emh5zjGLL0NgPGvuSbS9x7fT03dNGUadPIHtPmmU6JLnuZm8vT/fLlmU8U6CZjwNkJeINaJRruRPoUG6YREvxDGLbAZSLV9V+/nl0OisV1QS4JOv5vZKeWj+7BaEVguAHHJbzl9BdBROKl3I70AccF7VHYflB757MzVE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BjLCkf/Z; arc=none smtp.client-ip=209.85.128.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BjLCkf/Z" Received: by mail-wm1-f47.google.com with SMTP id 5b1f17b1804b1-4a0286c981dso31546515e9.2 for ; Tue, 06 Oct 2026 18:32:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1791336736; x=1791941536; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7KK+L3F3HJtHktm5mA74HhKfh2OLi0YG267n/zvJuk8=; b=BjLCkf/Z41SKV72EVIHFBvznWcyMPEWP7kDf9J0IRXKqA2wQx+3PEF8oz6QDATKwdL QnbjNaWDZGV6blmarS3pqz4o2Aac83hoM+Y4NoGzRgvbttmUudyWLC2VPLg8UhT8ltju E44bDy3DXq8bi/JVqk+IhyFdhafeGc2zgLOitlsbVXRwhRTlikS+C7qIXMrilrMjBXWX 7F7k5LP8VY4MBsecQUNnzfxKtoawG9gBvzbQK6mP0LMR/RSzs7m4uIly4JxXTWH7mR3p oyG9+peMkcDH1t4YTzoa5xaZjxmXVnVVU1dIyWdjIZRlBjLPCORw7J/Yln/NMwniYayM jlqw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791336736; x=1791941536; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7KK+L3F3HJtHktm5mA74HhKfh2OLi0YG267n/zvJuk8=; b=cSF/ZRJCo+iA51UD3DUADsj/xaGGi43KjxWRZdb9AaSt7BLMSAVBlw5GhjlhUv41Ej 8sHKuw8Y3vsdBckI7VkbZPC14kneOwMkir00fziscjej4V0lTUx7vEuvAHG7BGmJFIIR ICKbjjwIhrI6f3PxF+RkY3whlWPrMGV3ZSQ3qED4Di8T6Mv9UIEF5ALJ2rKOEpy75ScH mY3zJgZN4UAfTJ+dmsXxQOsdq+65SeNURIOWZjfE76IPrucCSlnyri9OuKIs7moYBB8P 93j0UPm2pIXDpbctNQA8hepmp2IPLqupzqLLPj8+/m/py5kxz2JWMTH4HjkZP9eP6qJb MGBA== X-Forwarded-Encrypted: i=1; AKwUvBwKBn1VV5J7bLFle60VRBiVoWvggctQixGeoiek7yHuBP64dMnbdncdub//Jflsq4u4gfQRFX++uw==@vger.kernel.org X-Gm-Message-State: AFuF++kz3pMsMM+zDXekq2ohLULZvS+DJlwIbfH2DyFloP5Kth2Enx2t dQAbYG+5DHHK3LPXIbrw2mmByuMQGVQKmxGhX1cYY+A8OeCty4QIPs4+ X-Gm-Gg: AYBFou3lyTM4S2R0ZtIlz5liELNqEyqeZ5LIqY7ZhThavclIoytiugKFVEwBJZMjm7h H8p393lLzo7cWKM9vHqH3vYgeQRhT9mhfcyRR5iLd5kA69RAdm04KYS8PCUVWZTUxkjHWx3f2Pe gsp/SUEYezGzWGsPLKg3enTijTghrALY2k+I858sFvU4C68Wi8GAoSomRzkSZPgjJhajIaczJZR M8BTIozuws4QfsO9TpMj6keLUWrWtrCssIr4b5BLLCPQ0TSYp13ADHaX6ut2RrIpJCyXgcIDoFM t3NjG1cHymiz76omEPKyyWrm59sMmfi6RiP3VxBnays084+saT6K3aDi/5AL4Id/WgRngpIBWjV FpufD1G2cezxnRjI0nWNkSMIbxlZULxJ9hTl1a34DSOIRevnkP1yGK/8+ODU5haUY8bFTzke+V2 4hrb+ACILcX03oIjTRkWUSulP6EfJRDqUz6LBuINm/Wy1q3no6CaFNoj7VKcb0ePAo8rpDhmjW2 vV21LH/f1BBIRKDCN0Na9qlEq9dVJmje2H9lAShwUuDN/ZXIXr/QhPiF9m31pDH/hYZuFJTl5IA wA== X-Received: by 2002:a05:600c:1551:b0:4a0:263c:704f with SMTP id 5b1f17b1804b1-4a1804436c7mr7142575e9.16.1791336736039; Tue, 06 Oct 2026 18:32:16 -0700 (PDT) Received: from 127.mynet ([2a01:4b00:bd21:4f00:7cc6:d3ca:494:116c]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4a17f6538b7sm21183615e9.13.2026.10.06.18.32.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 06 Oct 2026 18:32:14 -0700 (PDT) From: Pavel Begunkov To: linux-block@vger.kernel.org Cc: asml.silence@gmail.com, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, linux-nvme@lists.infradead.org, linux-fsdevel@vger.kernel.org, io-uring@vger.kernel.org, Christoph Hellwig , Sumit Semwal , =?UTF-8?q?Christian=20K=C3=B6nig?= , Keith Busch , Sagi Grimberg , Alexander Viro , Christian Brauner , Jan Kara , Andrew Morton , Jens Axboe , Nitesh Shetty , Kanchan Joshi , Anuj Gupta , Tushar Gohad , William Power , Matthew Brost , Alasdair Kergon , Mike Snitzer , Mikulas Patocka , Benjamin Marzinski , dm-devel@lists.linux.dev Subject: [PATCH v8 02/13] block: introduce dma map backed bio type Date: Wed, 7 Oct 2026 02:31:45 +0100 Message-ID: <17dc7954fc8632ca4fb80cd4c892fc0c5ce2d11b.1791336002.git.asml.silence@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: io-uring@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Premapped buffers don't require a generic bio_vec since these have already been dma mapped. Repurpose the bi_io_vec space to store dmabuf maps as they are mutually exclusive. The bio splitting differs from the normal path because it's already pre-mapped and for the block layer it's just an offset into the dma-buf. The actual segmentation is only available to the importer driver and not the block layer, however, it doesn't contain alignment gaps and we don't have to check it. For the same reason we can't precisely split by the number of segments, but we use ->min_seg_shift stored in the map to calculate the minimum number of bytes a request consisting of lim->max_segments full segments can cover and split by that. It's stricter and can add extra splitting. E.g. for a {4K, 4G} segmentation, the min segment size is 4K, and we'll split it into bios of (4K * lim->max_segments) bytes each, but it should be good enough for now to cover the most popular use cases. Suggested-by: Keith Busch Signed-off-by: Pavel Begunkov --- block/bio.c | 15 ++++++++++-- block/blk-merge.c | 50 +++++++++++++++++++++++++++++++++++++++ block/fops.c | 2 +- include/linux/bio.h | 9 +++---- include/linux/blk-mq.h | 7 ++++++ include/linux/blk_types.h | 14 ++++++++++- include/linux/bvec.h | 3 ++- 7 files changed, 91 insertions(+), 9 deletions(-) diff --git a/block/bio.c b/block/bio.c index b48091c7663f..1e0d9714c541 100644 --- a/block/bio.c +++ b/block/bio.c @@ -881,7 +881,11 @@ static int __bio_clone(struct bio *bio, struct bio *bio_src, gfp_t gfp) bio->bi_write_stream = bio_src->bi_write_stream; bio->bi_bvec_gap_bit = bio_src->bi_bvec_gap_bit; bio->bi_iter = bio_src->bi_iter; - bio->bi_io_vec = bio_src->bi_io_vec; + + if (op_is_dmabuf(bio->bi_opf)) + bio->bi_dmabuf_map = bio_src->bi_dmabuf_map; + else + bio->bi_io_vec = bio_src->bi_io_vec; if (bio->bi_bdev) { if (bio->bi_bdev == bio_src->bi_bdev && @@ -1204,16 +1208,23 @@ EXPORT_SYMBOL_GPL(__bio_release_pages); bool bio_iov_iter_set(struct bio *bio, const struct iov_iter *iter) { - if (!iov_iter_is_bvec(iter)) + if (!iov_iter_is_bvec(iter) && !iov_iter_is_dmabuf_map(iter)) return false; WARN_ON_ONCE(bio->bi_max_vecs); + static_assert(offsetof(struct bio, bi_io_vec) == + offsetof(struct bio, bi_dmabuf_map)); + static_assert(offsetof(struct iov_iter, bvec) == + offsetof(struct iov_iter, dmabuf_map)); + bio->bi_io_vec = (struct bio_vec *)iter->bvec; bio->bi_iter.bi_idx = 0; bio->bi_iter.bi_offset = iter->iov_offset; bio->bi_iter.bi_size = iov_iter_count(iter); bio_set_flag(bio, BIO_CLONED); + if (iov_iter_is_dmabuf_map(iter)) + bio->bi_opf |= REQ_NOMERGE | REQ_DMABUF; return true; } diff --git a/block/blk-merge.c b/block/blk-merge.c index 258a726071d1..18f014ff3314 100644 --- a/block/blk-merge.c +++ b/block/blk-merge.c @@ -9,6 +9,7 @@ #include #include #include +#include #include @@ -319,6 +320,41 @@ static inline unsigned int bvec_seg_gap(struct bio_vec *bvprv, return bv->bv_offset | (bvprv->bv_offset + bvprv->bv_len); } +static inline int bio_split_io_at_dmabuf(struct bio *bio, + const struct queue_limits *lim, unsigned *segs, + unsigned max_bytes, unsigned len_align_mask, + unsigned start_align_mask) +{ + unsigned bytes = min(bio->bi_iter.bi_size, max_bytes); + unsigned seg_shift = bio->bi_dmabuf_map->min_seg_shift; + unsigned offset = bio->bi_iter.bi_offset & ((1U << seg_shift) - 1); + u64 max_segs_bytes; + + /* + * dma-buf maps don't expose the underlying segmentation, but they're + * guaranteed to not have alignment gaps, we only need to check the + * start and length alignment. + */ + if ((bio->bi_iter.bi_offset & start_align_mask) || + (bio->bi_iter.bi_size & len_align_mask)) + return -EINVAL; + + /* Presented as a single contiguous range into the dma-buf */ + *segs = 1; + + /* + * Limit by the number of segments by using the minimal segment size. + * Any I/O consisting of N full segments should be able to cover at + * least N multiplied by the segment size. It's stricter than walking + * the segments and might cause extra splitting. + */ + max_segs_bytes = (u64)lim->max_segments << seg_shift; + bytes = min_t(u64, bytes, max_segs_bytes - offset); + if (bytes != bio->bi_iter.bi_size) + return bytes; + return 0; +} + /** * bio_split_io_at - check if and where to split a bio * @bio: [in] bio to be split @@ -346,6 +382,19 @@ int bio_split_io_at(struct bio *bio, const struct queue_limits *lim, len_align_mask |= (bc->bc_key->crypto_cfg.data_unit_size - 1); } + if (op_is_dmabuf(bio->bi_opf)) { + int ret; + + ret = bio_split_io_at_dmabuf(bio, lim, &nsegs, max_bytes, + len_align_mask, start_align_mask); + if (ret < 0) + return ret; + if (!ret) + goto out; + bytes = ret; + goto split; + } + bio_for_each_bvec(bv, bio, iter) { if (bv.bv_offset & start_align_mask || bv.bv_len & len_align_mask) @@ -376,6 +425,7 @@ int bio_split_io_at(struct bio *bio, const struct queue_limits *lim, bvprvp = &bvprv; } +out: *segs = nsegs; bio->bi_bvec_gap_bit = ffs(gaps); return 0; diff --git a/block/fops.c b/block/fops.c index 9905ed24a157..827dc9eecba7 100644 --- a/block/fops.c +++ b/block/fops.c @@ -362,7 +362,7 @@ static ssize_t __blkdev_direct_IO_async(struct kiocb *iocb, * Users don't rely on the iterator being in any particular * state for async I/O returning -EIOCBQUEUED, hence we can * avoid expensive iov_iter_advance(). Bypass - * bio_iov_iter_get_pages() and set the bvec directly. + * bio_iov_iter_get_pages() and set the bvec/dmabuf directly. */ if (!bio_iov_iter_set(bio, iter)) { ret = blkdev_iov_iter_get_pages(bio, iter, bdev); diff --git a/include/linux/bio.h b/include/linux/bio.h index 892ca469c570..7a794ce723b8 100644 --- a/include/linux/bio.h +++ b/include/linux/bio.h @@ -80,7 +80,8 @@ static inline bool bio_no_advance_iter(const struct bio *bio) { return bio_op(bio) == REQ_OP_DISCARD || bio_op(bio) == REQ_OP_SECURE_ERASE || - bio_op(bio) == REQ_OP_WRITE_ZEROES; + bio_op(bio) == REQ_OP_WRITE_ZEROES || + op_is_dmabuf(bio->bi_opf); } static inline void *bio_data(struct bio *bio) @@ -438,12 +439,12 @@ static inline void bio_wouldblock_error(struct bio *bio) /* * Calculate number of bvec segments that should be allocated to fit data - * pointed by @iter. If @iter is backed by bvec it's going to be reused - * instead of allocating a new one. + * pointed by @iter. If @iter is backed by a bvec or a dmabuf, the bvec array / + * the dma map are going to be reused, and so no extra allocation is required. */ static inline int bio_iov_vecs_to_alloc(struct iov_iter *iter, int max_segs) { - if (iov_iter_is_bvec(iter)) + if (iov_iter_is_bvec(iter) || iov_iter_is_dmabuf_map(iter)) return 0; return iov_iter_npages(iter, max_segs); } diff --git a/include/linux/blk-mq.h b/include/linux/blk-mq.h index af878597afb8..7c7504c84e09 100644 --- a/include/linux/blk-mq.h +++ b/include/linux/blk-mq.h @@ -1017,6 +1017,13 @@ static inline void *blk_mq_rq_to_pdu(struct request *rq) return rq + 1; } +static inline bool blk_mq_rq_is_dmabuf(struct request *rq) +{ + if (!IS_ENABLED(CONFIG_DMA_SHARED_BUFFER)) + return false; + return rq->bio && op_is_dmabuf(rq->bio->bi_opf); +} + static inline struct blk_mq_hw_ctx *queue_hctx(struct request_queue *q, int id) { struct blk_mq_hw_ctx *hctx; diff --git a/include/linux/blk_types.h b/include/linux/blk_types.h index 98e21b4cbf32..0cc09b975d8f 100644 --- a/include/linux/blk_types.h +++ b/include/linux/blk_types.h @@ -233,7 +233,12 @@ struct bio { atomic_t __bi_remaining; /* The actual vec list, preserved by bio_reset() */ - struct bio_vec *bi_io_vec; + union { + struct bio_vec *bi_io_vec; + /* Driver specific dma map, valid IFF REQ_DMABUF is set */ + struct dma_buf_io_map *bi_dmabuf_map; + }; + struct bvec_iter bi_iter; union { @@ -402,6 +407,7 @@ enum req_flag_bits { __REQ_DRV, /* for driver use */ __REQ_FS_PRIVATE, /* for file system (submitter) use */ __REQ_ATOMIC, /* for atomic write operations */ + __REQ_DMABUF, /* Using premmaped dma buffers */ /* * Command specific flags, keep last: */ @@ -434,6 +440,7 @@ enum req_flag_bits { #define REQ_DRV (__force blk_opf_t)(1ULL << __REQ_DRV) #define REQ_FS_PRIVATE (__force blk_opf_t)(1ULL << __REQ_FS_PRIVATE) #define REQ_ATOMIC (__force blk_opf_t)(1ULL << __REQ_ATOMIC) +#define REQ_DMABUF (__force blk_opf_t)(1ULL << __REQ_DMABUF) #define REQ_NOUNMAP (__force blk_opf_t)(1ULL << __REQ_NOUNMAP) @@ -487,6 +494,11 @@ static inline bool op_is_discard(blk_opf_t op) return (op & REQ_OP_MASK) == REQ_OP_DISCARD; } +static inline bool op_is_dmabuf(blk_opf_t op) +{ + return op & REQ_DMABUF; +} + /* * Check if a bio or request operation is a zone management operation. */ diff --git a/include/linux/bvec.h b/include/linux/bvec.h index fc566ee1c1ff..b63914ff56e3 100644 --- a/include/linux/bvec.h +++ b/include/linux/bvec.h @@ -108,7 +108,8 @@ struct bvec_iter { unsigned int bi_idx; /* - * Current offset in the bvec entry pointed to by `bi_idx`. + * Current offset in the bvec entry pointed to by `bi_idx` or into + * a dma-buf map. */ unsigned int bi_offset; } __packed __aligned(4); -- 2.54.0