[v3,16/27] btrfs: serialize data allocation and submit IOs

Message ID	20190808093038.4163421-17-naohiro.aota@wdc.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <linux-btrfs-owner@kernel.org> IronPort-SDR: 2TMmi3lXiElEHLzQzRKt/CNV7NfYcDOWjJAPMslmCZzzCPic+J62zgYt6kitjsGPCl14pg5M9r 571rQJslZMsobZTgfQLw5GdZ4ACz+rtdzqqTdI96KKIxQt8r8QCjX6nyM6ZK3fldwf3xnRM/Jy IgvKAXvZKj8bwUi56z5/rlfkS1ap5FEVUrfA0dwNAd+imKTDMjG6qn9kj0vwVRE+UdjsahGEpW nLWlIb1ddYwuB9Wha66rOZ3i7n5XzPUt5GPLy91F/Mh4tl+6+Ee6boHnUB78A9VPStJeMmLBLT gZ0= IronPort-SDR: G4BV1zFHLmJVmhe7m+kHNm2qdVYT0YJ1yXG/2IixIVwqm4OefdQE+qzX8XWZjhhwdgRuuf6gRc dMUPsJFH1Al1JpLmmO97s9bay1GY1ULVIoosIFWIVlRwsZdvh4RvKG4RPP5Z+x2uX/IP4B223t tU5uU2k+a11SJnR+QABtdihq0pj/Bv2Ny09ED6TE0r6BGO3a5PWnfNuaRNSiVW0OaRNAzVVg64 75Vm5wwS/6/iA+qq07E1u3h6yQJMFeDC9Nc41u5gJ/wlb+VHm07lMs1HfUBIOFLufhmkQiPqvi Mb8g+EnYUQiiT5RSfIxQv1SE IronPort-SDR: a0FXLMkufqMqShzalCB1sd/2Th8akOxgGvcLfVQcoOAq7L7Hz9J2bWToZW9F0boLELyNzPPK1l c3xjgyTwxtG+mxW7ytoeGax1HCE4aiNlYdIjVkIm6A+YJxouqR/ePDFm7ZLdqOl3oIVs0Loq1B 0/hmAooJn8nQBNERVdS4+krfLbQ7BIIcu1NRbnjo23yI3IVj8qdsaR4nmkkkSrxZuz0ZPcDSDW LK0Wz6jK1dCqSa7lBYQrn2PC7YqvW60WzlJ6lg7RV+uH49VzD6g37svBbgBOCvlSsCGwvpIFhY viY= From: Naohiro Aota <naohiro.aota@wdc.com> To: linux-btrfs@vger.kernel.org, David Sterba <dsterba@suse.com> Cc: Chris Mason <clm@fb.com>, Josef Bacik <josef@toxicpanda.com>, Nikolay Borisov <nborisov@suse.com>, Damien Le Moal <damien.lemoal@wdc.com>, Matias Bjorling <Matias.Bjorling@wdc.com>, Johannes Thumshirn <jthumshirn@suse.de>, Hannes Reinecke <hare@suse.com>, linux-fsdevel@vger.kernel.org, Naohiro Aota <naohiro.aota@wdc.com> Subject: [PATCH v3 16/27] btrfs: serialize data allocation and submit IOs Date: Thu, 8 Aug 2019 18:30:27 +0900 Message-Id: <20190808093038.4163421-17-naohiro.aota@wdc.com> In-Reply-To: <20190808093038.4163421-1-naohiro.aota@wdc.com> References: <20190808093038.4163421-1-naohiro.aota@wdc.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Sender: linux-btrfs-owner@vger.kernel.org Precedence: bulk
Series	btrfs zoned block device support \| expand [v3,00/27] btrfs zoned block device support [v3,01/27] btrfs: introduce HMZONED feature flag [v3,02/27] btrfs: Get zone information of zoned block devices [v3,03/27] btrfs: Check and enable HMZONED mode [v3,04/27] btrfs: disallow RAID5/6 in HMZONED mode [v3,05/27] btrfs: disallow space_cache in HMZONED mode [v3,06/27] btrfs: disallow NODATACOW in HMZONED mode [v3,07/27] btrfs: disable tree-log in HMZONED mode [v3,08/27] btrfs: disable fallocate in HMZONED mode [v3,09/27] btrfs: align device extent allocation to zone boundary [v3,10/27] btrfs: do sequential extent allocation in HMZONED mode [v3,11/27] btrfs: make unmirroed BGs readonly only if we have at least one writable BG [v3,12/27] btrfs: ensure metadata space available on/after degraded mount in HMZONED [v3,13/27] btrfs: reset zones of unused block groups [v3,14/27] btrfs: limit super block locations in HMZONED mode [v3,15/27] btrfs: redirty released extent buffers in sequential BGs [v3,16/27] btrfs: serialize data allocation and submit IOs [v3,17/27] btrfs: implement atomic compressed IO submission [v3,18/27] btrfs: support direct write IO in HMZONED [v3,19/27] btrfs: serialize meta IOs on HMZONED mode [v3,20/27] btrfs: wait existing extents before truncating [v3,21/27] btrfs: avoid async checksum/submit on HMZONED mode [v3,22/27] btrfs: disallow mixed-bg in HMZONED mode [v3,23/27] btrfs: disallow inode_cache in HMZONED mode [v3,24/27] btrfs: support dev-replace in HMZONED mode [v3,25/27] btrfs: enable relocation in HMZONED mode [v3,26/27] btrfs: relocate block group to repair IO failure in HMZONED [v3,27/27] btrfs: enable to mount HMZONED incompat flag

Message ID

20190808093038.4163421-17-naohiro.aota@wdc.com (mailing list archive)

State

New, archived

Headers

IronPort-SDR: 
 2TMmi3lXiElEHLzQzRKt/CNV7NfYcDOWjJAPMslmCZzzCPic+J62zgYt6kitjsGPCl14pg5M9r
 571rQJslZMsobZTgfQLw5GdZ4ACz+rtdzqqTdI96KKIxQt8r8QCjX6nyM6ZK3fldwf3xnRM/Jy
 IgvKAXvZKj8bwUi56z5/rlfkS1ap5FEVUrfA0dwNAd+imKTDMjG6qn9kj0vwVRE+UdjsahGEpW
 nLWlIb1ddYwuB9Wha66rOZ3i7n5XzPUt5GPLy91F/Mh4tl+6+Ee6boHnUB78A9VPStJeMmLBLT
 gZ0=
IronPort-SDR: 
 G4BV1zFHLmJVmhe7m+kHNm2qdVYT0YJ1yXG/2IixIVwqm4OefdQE+qzX8XWZjhhwdgRuuf6gRc
 dMUPsJFH1Al1JpLmmO97s9bay1GY1ULVIoosIFWIVlRwsZdvh4RvKG4RPP5Z+x2uX/IP4B223t
 tU5uU2k+a11SJnR+QABtdihq0pj/Bv2Ny09ED6TE0r6BGO3a5PWnfNuaRNSiVW0OaRNAzVVg64
 75Vm5wwS/6/iA+qq07E1u3h6yQJMFeDC9Nc41u5gJ/wlb+VHm07lMs1HfUBIOFLufhmkQiPqvi
 Mb8g+EnYUQiiT5RSfIxQv1SE
IronPort-SDR: 
 a0FXLMkufqMqShzalCB1sd/2Th8akOxgGvcLfVQcoOAq7L7Hz9J2bWToZW9F0boLELyNzPPK1l
 c3xjgyTwxtG+mxW7ytoeGax1HCE4aiNlYdIjVkIm6A+YJxouqR/ePDFm7ZLdqOl3oIVs0Loq1B
 0/hmAooJn8nQBNERVdS4+krfLbQ7BIIcu1NRbnjo23yI3IVj8qdsaR4nmkkkSrxZuz0ZPcDSDW
 LK0Wz6jK1dCqSa7lBYQrn2PC7YqvW60WzlJ6lg7RV+uH49VzD6g37svBbgBOCvlSsCGwvpIFhY
 viY=
From: Naohiro Aota <naohiro.aota@wdc.com>
To: linux-btrfs@vger.kernel.org, David Sterba <dsterba@suse.com>
Cc: Chris Mason <clm@fb.com>, Josef Bacik <josef@toxicpanda.com>,
        Nikolay Borisov <nborisov@suse.com>,
        Damien Le Moal <damien.lemoal@wdc.com>,
        Matias Bjorling <Matias.Bjorling@wdc.com>,
        Johannes Thumshirn <jthumshirn@suse.de>,
        Hannes Reinecke <hare@suse.com>, linux-fsdevel@vger.kernel.org,
        Naohiro Aota <naohiro.aota@wdc.com>
Subject: [PATCH v3 16/27] btrfs: serialize data allocation and submit IOs
Date: Thu,  8 Aug 2019 18:30:27 +0900
Message-Id: <20190808093038.4163421-17-naohiro.aota@wdc.com>
In-Reply-To: <20190808093038.4163421-1-naohiro.aota@wdc.com>
References: <20190808093038.4163421-1-naohiro.aota@wdc.com>
MIME-Version: 1.0
Content-Transfer-Encoding: 8bit
Sender: linux-btrfs-owner@vger.kernel.org
Precedence: bulk

Series

btrfs zoned block device support | expand

Commit Message

Naohiro Aota Aug. 8, 2019, 9:30 a.m. UTC

To preserve sequential write pattern on the drives, we must serialize
allocation and submit_bio. This commit add per-block group mutex
"zone_io_lock" and find_free_extent_seq() hold the lock. The lock is kept
even after returning from find_free_extent(). It is released when submiting
IOs corresponding to the allocation is completed.

Implementing such behavior under __extent_writepage_io is almost impossible
because once pages are unlocked we are not sure when submiting IOs for an
allocated region is finished or not. Instead, this commit add
run_delalloc_hmzoned() to write out non-compressed data IOs at once using
extent_write_locked_rage(). After the write, we can call
btrfs_hmzoned_unlock_allocation() to unlock the block group for new
allocation.

Signed-off-by: Naohiro Aota <naohiro.aota@wdc.com>
---
 fs/btrfs/ctree.h       |  1 +
 fs/btrfs/extent-tree.c |  5 +++++
 fs/btrfs/hmzoned.h     | 34 +++++++++++++++++++++++++++++++
 fs/btrfs/inode.c       | 45 ++++++++++++++++++++++++++++++++++++++++--
 4 files changed, 83 insertions(+), 2 deletions(-)

diff --git a/fs/btrfs/ctree.h b/fs/btrfs/ctree.h
index 3d31a1960c4d..1e924c0d1210 100644
--- a/fs/btrfs/ctree.h
+++ b/fs/btrfs/ctree.h
@@ -619,6 +619,7 @@  struct btrfs_block_group_cache {
 	 * zone.
 	 */
 	u64 alloc_offset;
+	struct mutex zone_io_lock;
 };
 
 /* delayed seq elem */
diff --git a/fs/btrfs/extent-tree.c b/fs/btrfs/extent-tree.c
index bc95a73a762d..5b1a9e607555 100644
--- a/fs/btrfs/extent-tree.c
+++ b/fs/btrfs/extent-tree.c
@@ -5532,6 +5532,7 @@  static int find_free_extent_seq(struct btrfs_block_group_cache *cache,
 	if (cache->alloc_type != BTRFS_ALLOC_SEQ)
 		return 1;
 
+	btrfs_hmzoned_data_io_lock(cache);
 	spin_lock(&space_info->lock);
 	spin_lock(&cache->lock);
 
@@ -5563,6 +5564,9 @@  static int find_free_extent_seq(struct btrfs_block_group_cache *cache,
 out:
 	spin_unlock(&cache->lock);
 	spin_unlock(&space_info->lock);
+	/* if succeeds, unlock after submit_bio */
+	if (ret)
+		btrfs_hmzoned_data_io_unlock(cache);
 	return ret;
 }
 
@@ -8104,6 +8108,7 @@  btrfs_create_block_group_cache(struct btrfs_fs_info *fs_info,
 	btrfs_init_free_space_ctl(cache);
 	atomic_set(&cache->trimming, 0);
 	mutex_init(&cache->free_space_lock);
+	mutex_init(&cache->zone_io_lock);
 	btrfs_init_full_stripe_locks_tree(&cache->full_stripe_locks_root);
 	cache->alloc_type = BTRFS_ALLOC_FIT;
 
diff --git a/fs/btrfs/hmzoned.h b/fs/btrfs/hmzoned.h
index 3a73c3c5e1da..a8e7286708d4 100644
--- a/fs/btrfs/hmzoned.h
+++ b/fs/btrfs/hmzoned.h
@@ -39,6 +39,7 @@  int btrfs_hmzoned_check_metadata_space(struct btrfs_fs_info *fs_info);
 void btrfs_redirty_list_add(struct btrfs_transaction *trans,
 			    struct extent_buffer *eb);
 void btrfs_free_redirty_list(struct btrfs_transaction *trans);
+void btrfs_hmzoned_data_io_unlock_at(struct inode *inode, u64 start, u64 len);
 
 static inline bool btrfs_dev_is_sequential(struct btrfs_device *device, u64 pos)
 {
@@ -140,4 +141,37 @@  static inline bool btrfs_check_super_location(struct btrfs_device *device,
 		!btrfs_dev_is_sequential(device, pos);
 }
 
+
+static inline void btrfs_hmzoned_data_io_lock(
+	struct btrfs_block_group_cache *cache)
+{
+	/* No need to lock metadata BGs or non-sequential BGs */
+	if (!(cache->flags & BTRFS_BLOCK_GROUP_DATA) ||
+	    cache->alloc_type != BTRFS_ALLOC_SEQ)
+		return;
+	mutex_lock(&cache->zone_io_lock);
+}
+
+static inline void btrfs_hmzoned_data_io_unlock(
+	struct btrfs_block_group_cache *cache)
+{
+	if (!(cache->flags & BTRFS_BLOCK_GROUP_DATA) ||
+	    cache->alloc_type != BTRFS_ALLOC_SEQ)
+		return;
+	mutex_unlock(&cache->zone_io_lock);
+}
+
+static inline void btrfs_hmzoned_data_io_unlock_logical(
+	struct btrfs_fs_info *fs_info, u64 logical)
+{
+	struct btrfs_block_group_cache *cache;
+
+	if (!btrfs_fs_incompat(fs_info, HMZONED))
+		return;
+
+	cache = btrfs_lookup_block_group(fs_info, logical);
+	btrfs_hmzoned_data_io_unlock(cache);
+	btrfs_put_block_group(cache);
+}
+
 #endif
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index ee582a36653d..d504200c9767 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -48,6 +48,7 @@ 
 #include "qgroup.h"
 #include "dedupe.h"
 #include "delalloc-space.h"
+#include "hmzoned.h"
 
 struct btrfs_iget_args {
 	struct btrfs_key *location;
@@ -1279,6 +1280,39 @@  static int cow_file_range_async(struct inode *inode, struct page *locked_page,
 	return 0;
 }
 
+static noinline int run_delalloc_hmzoned(struct inode *inode,
+					 struct page *locked_page, u64 start,
+					 u64 end, int *page_started,
+					 unsigned long *nr_written)
+{
+	struct extent_map *em;
+	u64 logical;
+	int ret;
+
+	ret = cow_file_range(inode, locked_page, start, end,
+			     end, page_started, nr_written, 0, NULL);
+	if (ret)
+		return ret;
+
+	if (*page_started)
+		return 0;
+
+	em = btrfs_get_extent(BTRFS_I(inode), NULL, 0, start, end - start + 1,
+			      0);
+	ASSERT(em != NULL && em->block_start < EXTENT_MAP_LAST_BYTE);
+	logical = em->block_start;
+	free_extent_map(em);
+
+	__set_page_dirty_nobuffers(locked_page);
+	account_page_redirty(locked_page);
+	extent_write_locked_range(inode, start, end, WB_SYNC_ALL);
+	*page_started = 1;
+
+	btrfs_hmzoned_data_io_unlock_logical(btrfs_sb(inode->i_sb), logical);
+
+	return 0;
+}
+
 static noinline int csum_exist_in_range(struct btrfs_fs_info *fs_info,
 					u64 bytenr, u64 num_bytes)
 {
@@ -1645,17 +1679,24 @@  int btrfs_run_delalloc_range(struct inode *inode, struct page *locked_page,
 	int ret;
 	int force_cow = need_force_cow(inode, start, end);
 	unsigned int write_flags = wbc_to_write_flags(wbc);
+	int do_compress = inode_can_compress(inode) &&
+		inode_need_compress(inode, start, end);
+	int hmzoned = btrfs_fs_incompat(btrfs_sb(inode->i_sb), HMZONED);
 
 	if (BTRFS_I(inode)->flags & BTRFS_INODE_NODATACOW && !force_cow) {
+		ASSERT(!hmzoned);
 		ret = run_delalloc_nocow(inode, locked_page, start, end,
 					 page_started, 1, nr_written);
 	} else if (BTRFS_I(inode)->flags & BTRFS_INODE_PREALLOC && !force_cow) {
+		ASSERT(!hmzoned);
 		ret = run_delalloc_nocow(inode, locked_page, start, end,
 					 page_started, 0, nr_written);
-	} else if (!inode_can_compress(inode) ||
-		   !inode_need_compress(inode, start, end)) {
+	} else if (!do_compress && !hmzoned) {
 		ret = cow_file_range(inode, locked_page, start, end, end,
 				      page_started, nr_written, 1, NULL);
+	} else if (!do_compress && hmzoned) {
+		ret = run_delalloc_hmzoned(inode, locked_page, start, end,
+					   page_started, nr_written);
 	} else {
 		set_bit(BTRFS_INODE_HAS_ASYNC_EXTENT,
 			&BTRFS_I(inode)->runtime_flags);

[v3,16/27] btrfs: serialize data allocation and submit IOs

Commit Message

Patch