[v4,1/2] iomap: fix zero padding data issue in concurrent append writes

During concurrent append writes to XFS filesystem, zero padding data
may appear in the file after power failure. This happens due to imprecise
disk size updates when handling write completion.

Consider this scenario with concurrent append writes same file:

  Thread 1:                  Thread 2:
  ------------               -----------
  write [A, A+B]
  update inode size to A+B
  submit I/O [A, A+BS]
                             write [A+B, A+B+C]
                             update inode size to A+B+C
  <I/O completes, updates disk size to min(A+B+C, A+BS)>
  <power failure>

After reboot:
  1) with A+B+C < A+BS, the file has zero padding in range [A+B, A+B+C]

  |<         Block Size (BS)      >|
  |DDDDDDDDDDDDDDDD0000000000000000|
  ^               ^        ^
  A              A+B     A+B+C
                         (EOF)

  2) with A+B+C > A+BS, the file has zero padding in range [A+B, A+BS]

  |<         Block Size (BS)      >|<           Block Size (BS)    >|
  |DDDDDDDDDDDDDDDD0000000000000000|00000000000000000000000000000000|
  ^               ^                ^               ^
  A              A+B              A+BS           A+B+C
                                  (EOF)

  D = Valid Data
  0 = Zero Padding

The issue stems from disk size being set to min(io_offset + io_size,
inode->i_size) at I/O completion. Since io_offset+io_size is block
size granularity, it may exceed the actual valid file data size. In
the case of concurrent append writes, inode->i_size may be larger
than the actual range of valid file data written to disk, leading to
inaccurate disk size updates.

This patch changes the meaning of io_size to represent the size of
valid data within eof in ioend, while the extent size of ioend can
be obtained by rounding up to block size in wrapper function.
This function is specifically used for ioend grow/merge management:
1. In concurrent writes, when one write's io_size is truncated due
   to non-block-aligned file size while another write extends the file
   size, if these two writes are physically and logically contiguous
   at block boundaries, rounding up io_size to block boundaries helps
   grow the first write's ioend and share this ioend between both
   writes.
2. During IO completion, we try to merge physically and logically
   contiguous ioends before completion to minimize the number of
   transactions. Rounding up io_size to block boundaries helps merge
   ioends whose io_size is not block-aligned.

Another benefit is that it makes the xfs_ioend_is_append() check more
accurate, which can reduce unnecessary end bio callbacks of xfs_end_bio()
in certain scenarios, such as repeated writes at the file tail without
extending the file size.

Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Signed-off-by: Long Li <leo.lilong@huawei.com>
Reviewed-by: Brian Foster <bfoster@redhat.com>
---
v3->v4:
  1. collect reviewed tag
  2. Modify the comment of io_size and iomap_ioend_size_aligned().
  3. Add explain of iomap_ioend_size_aligned() to commit message.
 fs/iomap/buffered-io.c | 37 +++++++++++++++++++++++++++++++------
 include/linux/iomap.h  |  2 +-
 2 files changed, 32 insertions(+), 7 deletions(-)

Message ID	20241125023341.2816630-1-leo.lilong@huawei.com (mailing list archive)
State	New
Headers	show Received: from szxga07-in.huawei.com (szxga07-in.huawei.com [45.249.212.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 06D302F2D; Mon, 25 Nov 2024 02:55:46 +0000 (UTC) From: Long Li <leo.lilong@huawei.com> To: <brauner@kernel.org>, <djwong@kernel.org>, <cem@kernel.org> CC: <linux-xfs@vger.kernel.org>, <linux-fsdevel@vger.kernel.org>, <yi.zhang@huawei.com>, <houtao1@huawei.com>, <leo.lilong@huawei.com>, <yangerkun@huawei.com> Subject: [PATCH v4 1/2] iomap: fix zero padding data issue in concurrent append writes Date: Mon, 25 Nov 2024 10:33:40 +0800 Message-ID: <20241125023341.2816630-1-leo.lilong@huawei.com> Precedence: bulk MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Content-Type: text/plain
Series	[v4,1/2] iomap: fix zero padding data issue in concurrent append writes \| expand [v4,1/2] iomap: fix zero padding data issue in concurrent append writes [v4,2/2] xfs: clean up xfs_end_ioend() to reuse local variables

[v4,1/2] iomap: fix zero padding data issue in concurrent append writes

Commit Message

Comments

Patch