[5.15] btrfs: fix space cache corruption and potential double allocations

From: Omar Sandoval <osandov@fb.com>

From: Omar Sandoval <osandov@fb.com>

commit ced8ecf026fd8084cf175530ff85c76d6085d715 upstream.

When testing space_cache v2 on a large set of machines, we encountered a
few symptoms:

1. "unable to add free space :-17" (EEXIST) errors.
2. Missing free space info items, sometimes caught with a "missing free
   space info for X" error.
3. Double-accounted space: ranges that were allocated in the extent tree
   and also marked as free in the free space tree, ranges that were
   marked as allocated twice in the extent tree, or ranges that were
   marked as free twice in the free space tree. If the latter made it
   onto disk, the next reboot would hit the BUG_ON() in
   add_new_free_space().
4. On some hosts with no on-disk corruption or error messages, the
   in-memory space cache (dumped with drgn) disagreed with the free
   space tree.

All of these symptoms have the same underlying cause: a race between
caching the free space for a block group and returning free space to the
in-memory space cache for pinned extents causes us to double-add a free
range to the space cache. This race exists when free space is cached
from the free space tree (space_cache=v2) or the extent tree
(nospace_cache, or space_cache=v1 if the cache needs to be regenerated).
struct btrfs_block_group::last_byte_to_unpin and struct
btrfs_block_group::progress are supposed to protect against this race,
but commit d0c2f4fa555e ("btrfs: make concurrent fsyncs wait less when
waiting for a transaction commit") subtly broke this by allowing
multiple transactions to be unpinning extents at the same time.

Specifically, the race is as follows:

1. An extent is deleted from an uncached block group in transaction A.
2. btrfs_commit_transaction() is called for transaction A.
3. btrfs_run_delayed_refs() -> __btrfs_free_extent() runs the delayed
   ref for the deleted extent.
4. __btrfs_free_extent() -> do_free_extent_accounting() ->
   add_to_free_space_tree() adds the deleted extent back to the free
   space tree.
5. do_free_extent_accounting() -> btrfs_update_block_group() ->
   btrfs_cache_block_group() queues up the block group to get cached.
   block_group->progress is set to block_group->start.
6. btrfs_commit_transaction() for transaction A calls
   switch_commit_roots(). It sets block_group->last_byte_to_unpin to
   block_group->progress, which is block_group->start because the block
   group hasn't been cached yet.
7. The caching thread gets to our block group. Since the commit roots
   were already switched, load_free_space_tree() sees the deleted extent
   as free and adds it to the space cache. It finishes caching and sets
   block_group->progress to U64_MAX.
8. btrfs_commit_transaction() advances transaction A to
   TRANS_STATE_SUPER_COMMITTED.
9. fsync calls btrfs_commit_transaction() for transaction B. Since
   transaction A is already in TRANS_STATE_SUPER_COMMITTED and the
   commit is for fsync, it advances.
10. btrfs_commit_transaction() for transaction B calls
    switch_commit_roots(). This time, the block group has already been
    cached, so it sets block_group->last_byte_to_unpin to U64_MAX.
11. btrfs_commit_transaction() for transaction A calls
    btrfs_finish_extent_commit(), which calls unpin_extent_range() for
    the deleted extent. It sees last_byte_to_unpin set to U64_MAX (by
    transaction B!), so it adds the deleted extent to the space cache
    again!

This explains all of our symptoms above:

* If the sequence of events is exactly as described above, when the free
  space is re-added in step 11, it will fail with EEXIST.
* If another thread reallocates the deleted extent in between steps 7
  and 11, then step 11 will silently re-add that space to the space
  cache as free even though it is actually allocated. Then, if that
  space is allocated *again*, the free space tree will be corrupted
  (namely, the wrong item will be deleted).
* If we don't catch this free space tree corruption, it will continue
  to get worse as extents are deleted and reallocated.

The v1 space_cache is synchronously loaded when an extent is deleted
(btrfs_update_block_group() with alloc=0 calls btrfs_cache_block_group()
with load_cache_only=1), so it is not normally affected by this bug.
However, as noted above, if we fail to load the space cache, we will
fall back to caching from the extent tree and may hit this bug.

The easiest fix for this race is to also make caching from the free
space tree or extent tree synchronous. Josef tested this and found no
performance regressions.

A few extra changes fall out of this change. Namely, this fix does the
following, with step 2 being the crucial fix:

1. Factor btrfs_caching_ctl_wait_done() out of
   btrfs_wait_block_group_cache_done() to allow waiting on a caching_ctl
   that we already hold a reference to.
2. Change the call in btrfs_cache_block_group() of
   btrfs_wait_space_cache_v1_finished() to
   btrfs_caching_ctl_wait_done(), which makes us wait regardless of the
   space_cache option.
3. Delete the now unused btrfs_wait_space_cache_v1_finished() and
   space_cache_v1_done().
4. Change btrfs_cache_block_group()'s `int load_cache_only` parameter to
   `bool wait` to more accurately describe its new meaning.
5. Change a few callers which had a separate call to
   btrfs_wait_block_group_cache_done() to use wait = true instead.
6. Make btrfs_wait_block_group_cache_done() static now that it's not
   used outside of block-group.c anymore.

Fixes: d0c2f4fa555e ("btrfs: make concurrent fsyncs wait less when waiting for a transaction commit")
CC: stable@vger.kernel.org # 5.12+
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Omar Sandoval <osandov@fb.com>
Signed-off-by: David Sterba <dsterba@suse.com>
---
Hi,

This is the backport of commit ced8ecf026fd8084cf175530ff85c76d6085d715
to the 5.15 stable branch. Please consider it for the next 5.15 stable
release.

Thanks,
Omar

 fs/btrfs/block-group.c | 47 ++++++++++++++----------------------------
 fs/btrfs/block-group.h |  4 +---
 fs/btrfs/ctree.h       |  1 -
 fs/btrfs/extent-tree.c | 30 ++++++---------------------
 4 files changed, 22 insertions(+), 60 deletions(-)

Message ID	b07b2c2bcef831de5eaa6d2e61d46b92db4d41d5.1662055214.git.osandov@fb.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <linux-btrfs-owner@kernel.org> X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 21416ECAAD3 for <linux-btrfs@archiver.kernel.org>; Thu, 1 Sep 2022 18:03:48 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S232578AbiIASDq (ORCPT <rfc822;linux-btrfs@archiver.kernel.org>); Thu, 1 Sep 2022 14:03:46 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:46680 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S232151AbiIASDp (ORCPT <rfc822;linux-btrfs@vger.kernel.org>); Thu, 1 Sep 2022 14:03:45 -0400 Received: from mail-pf1-x42e.google.com (mail-pf1-x42e.google.com [IPv6:2607:f8b0:4864:20::42e]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 02D886F275 for <linux-btrfs@vger.kernel.org>; Thu, 1 Sep 2022 11:03:44 -0700 (PDT) Received: by mail-pf1-x42e.google.com with SMTP id y127so18217634pfy.5 for <linux-btrfs@vger.kernel.org>; Thu, 01 Sep 2022 11:03:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=osandov-com.20210112.gappssmtp.com; s=20210112; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date; bh=/Kb0s03L08yISPG2LWL2pe0LZ3W0qJJz4MBqRAx8pUc=; b=mMJR2Zg33m7YBHSWJIN0C5EslNb3T06OeJUp2xpy5kbI/APxW1ukPeci1sbRQxQth3 DVpgKsTtradf1A/ugFQwcl7IzZxI/vn1ZHX8ZxdfDAZKC/8Djfzjd/4l43jji5tXFPNs IEmE/yi1Q2qNBUnTkiiw5TmMmvQ1T+KyaYYz91BhYo0otRDX4qjkRUKbR51JU3+El3zt lYuyPXcE2HnXNCaVGxCE1R/sEwNU5BucfNCk/Sk6OoJAzeyBPu9cdFeI3UpS9YNLMb2u 39HvHQm4q02wvRwFQovDqModjhlCKrqVYtwUPEIDr7rp6KX5eRLv+TOdGv8rHgFtFBAm DK5g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date; bh=/Kb0s03L08yISPG2LWL2pe0LZ3W0qJJz4MBqRAx8pUc=; b=A1VqlDKBUQaNmCaEPbA6gbJv3OXv7MtWnPMQdVV/SyNbbRCBaYMUMVqZ634b+qlBwN sGzhJKpcSYNRiACwzCrSx2gFmkMO01Smoy5vJeJ7ScWBAajo8PoBFcVk35/5Av67pbKR ITpxSX3XT6E6Sujn7a8E3Dv8IOYeuEA++skarAZwumt8szyLTGizNSiOc9EifiSyTKsS bPGcKCRZ87KEbBr8aESjBiUg9M5nflGM5H6kK0OOe3V0hAAhZotirNDzXauURV52JDNn 7pW/RzUJyV2SIsk0nVqP7PPPaPQysI1b9+RQeLg3RvrMHTjc4wsmzbkXzkdZYbp7tlrK QDNg== X-Gm-Message-State: ACgBeo1DlPJKybqgZI2uCnIkxTM6k5IcSw3gWHCQIsrYdth1TtTb+Ld/ WqIfMLjzC8gRAwGEMctWavMmUWZoayiptw== X-Google-Smtp-Source: AA6agR7Hphykh1RJEueNKnnRn3QwV7V4P4ZJCIyWppRGXSDB16y94TBTrXQjh+CIVwAR5VRSFp4evw== X-Received: by 2002:a05:6a00:140d:b0:52a:d561:d991 with SMTP id l13-20020a056a00140d00b0052ad561d991mr32378925pfu.46.1662055422909; Thu, 01 Sep 2022 11:03:42 -0700 (PDT) Received: from relinquished.thefacebook.com ([2620:10d:c090:500::594d]) by smtp.gmail.com with ESMTPSA id q5-20020aa79605000000b00536b8f91806sm13614706pfg.198.2022.09.01.11.03.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 01 Sep 2022 11:03:41 -0700 (PDT) From: Omar Sandoval <osandov@osandov.com> To: linux-btrfs@vger.kernel.org, stable@vger.kernel.org Cc: kernel-team@fb.com Subject: [PATCH 5.15] btrfs: fix space cache corruption and potential double allocations Date: Thu, 1 Sep 2022 11:03:35 -0700 Message-Id: <b07b2c2bcef831de5eaa6d2e61d46b92db4d41d5.1662055214.git.osandov@fb.com> X-Mailer: git-send-email 2.37.3 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Precedence: bulk List-ID: <linux-btrfs.vger.kernel.org> X-Mailing-List: linux-btrfs@vger.kernel.org
Series	[5.15] btrfs: fix space cache corruption and potential double allocations \| expand [5.15] btrfs: fix space cache corruption and potential double allocations

[5.15] btrfs: fix space cache corruption and potential double allocations

Commit Message

Comments

Patch