From patchwork Tue Jun 25 11:44:14 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: "Pankaj Raghav (Samsung)" X-Patchwork-Id: 13710923 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 7053AC2BBCA for ; Tue, 25 Jun 2024 11:44:47 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id ED4056B02EE; Tue, 25 Jun 2024 07:44:46 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id E83F86B02EF; Tue, 25 Jun 2024 07:44:46 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id CFEDC6B02F0; Tue, 25 Jun 2024 07:44:46 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0017.hostedemail.com [216.40.44.17]) by kanga.kvack.org (Postfix) with ESMTP id ABAAE6B02EE for ; Tue, 25 Jun 2024 07:44:46 -0400 (EDT) Received: from smtpin14.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay08.hostedemail.com (Postfix) with ESMTP id 242A01414D1 for ; Tue, 25 Jun 2024 11:44:46 +0000 (UTC) X-FDA: 82269228972.14.2792A07 Received: from mout-p-101.mailbox.org (mout-p-101.mailbox.org [80.241.56.151]) by imf08.hostedemail.com (Postfix) with ESMTP id 5099A160018 for ; Tue, 25 Jun 2024 11:44:44 +0000 (UTC) Authentication-Results: imf08.hostedemail.com; dkim=pass header.d=pankajraghav.com header.s=MBO0001 header.b=PiG+mKE7; spf=pass (imf08.hostedemail.com: domain of kernel@pankajraghav.com designates 80.241.56.151 as permitted sender) smtp.mailfrom=kernel@pankajraghav.com; dmarc=pass (policy=quarantine) header.from=pankajraghav.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1719315877; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=N1ycwWKGYSoq9aKuREaT9SOLOX0Nv8uK+1dIKxx6K5c=; b=UM6uHFjWSR3nLAz+bceyHZwTBefhumxGTH0mqO23y1EdvhsGuyusEcqN+G1JKI6tev8LG9 Pugwy+CaWYGKvZJ7LHMHAHIzMHLkHNtorRmcZ3/FwjNrVeSH6HxuUsX2Fd1kx8F5XepcH2 ZFYR2fgu/K85b2Q/ScDk2Z1vqlY32u4= ARC-Authentication-Results: i=1; imf08.hostedemail.com; dkim=pass header.d=pankajraghav.com header.s=MBO0001 header.b=PiG+mKE7; spf=pass (imf08.hostedemail.com: domain of kernel@pankajraghav.com designates 80.241.56.151 as permitted sender) smtp.mailfrom=kernel@pankajraghav.com; dmarc=pass (policy=quarantine) header.from=pankajraghav.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1719315877; a=rsa-sha256; cv=none; b=og2uRV4HknRoopcd9KPH3MbdI8GTD3JrslvHKyg94qJDlKYANAbwSPyNk8eTxDiapij+gY MnZ+paza1knMSbxqi7QFdqYJXlZUtt47t8/49+Jg0PZgDlYgzJdp3a9GsUE+EQCWmSd9Nf EFG6xWwI1x2F+S6W+V53nOs1vKjZIkA= Received: from smtp2.mailbox.org (smtp2.mailbox.org [10.196.197.2]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (4096 bits) server-digest SHA256) (No client certificate requested) by mout-p-101.mailbox.org (Postfix) with ESMTPS id 4W7jfx1tRlz9sYX; Tue, 25 Jun 2024 13:44:41 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=pankajraghav.com; s=MBO0001; t=1719315881; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=N1ycwWKGYSoq9aKuREaT9SOLOX0Nv8uK+1dIKxx6K5c=; b=PiG+mKE7HVDbyP+6P1IA+X3BgQEg8VFBTpkSS4qWE7AZ6/0YDNba7z63Q9TR7q5pMnDCHz 1YeFOEzjf0Aa4Lba+wO3gLyBTd5mQYepm7+RAv1OkHLr1TuSoHheqCd4EsqR6mL5AtkTjz 5OUQKA27sBsfIDLrwyuhdqWw8f6uL7NK+N/AVMWnL7C/wZBZjecDlbelI3PsgLHqkADWIq S1Pu8YnXEr6jWBqXn0I5nIdgnYbI5W5o18+JwqEiF05vdLSgfqmHUv6sQftCJPn6/au2DO 3MJkCFwSJhu8unImJO0t+ehOs6N9SeIE0q639VqHhhfLnG9l8cQ4GcTfGxcYVg== From: "Pankaj Raghav (Samsung)" To: david@fromorbit.com, willy@infradead.org, chandan.babu@oracle.com, djwong@kernel.org, brauner@kernel.org, akpm@linux-foundation.org Cc: linux-kernel@vger.kernel.org, yang@os.amperecomputing.com, linux-mm@kvack.org, john.g.garry@oracle.com, linux-fsdevel@vger.kernel.org, hare@suse.de, p.raghav@samsung.com, mcgrof@kernel.org, gost.dev@samsung.com, cl@os.amperecomputing.com, linux-xfs@vger.kernel.org, kernel@pankajraghav.com, hch@lst.de, Zi Yan Subject: [PATCH v8 04/10] mm: split a folio in minimum folio order chunks Date: Tue, 25 Jun 2024 11:44:14 +0000 Message-ID: <20240625114420.719014-5-kernel@pankajraghav.com> In-Reply-To: <20240625114420.719014-1-kernel@pankajraghav.com> References: <20240625114420.719014-1-kernel@pankajraghav.com> MIME-Version: 1.0 X-Rspam-User: X-Rspamd-Server: rspam01 X-Rspamd-Queue-Id: 5099A160018 X-Stat-Signature: ogw1io5k58d9yj6wz8s1gf3y64bxtiyu X-HE-Tag: 1719315884-464540 X-HE-Meta: U2FsdGVkX19hc54ysxSEAqGFgTV0ZwmevdUlRu9Vz/BW5+1WPdepC3GSKjwNhDvGBdsIFi3dOJ4TK2RxzurRKoifnRfBLZ2yg5L51Es8v1OjuTS3FDsWtg/9B0t3uX96HFOBabqrT7NxSx9JmZf84ViSCt36Wj2uu669kJ9jacqKPrOnwivHwyJrjfhqzI4DWEE5+IZN5UGUb2iJ/AGve+B6ZclEapCcpdMfAgcHkTBZylhe+45km7OnvEVvD8jX8gvHtON5aOsS53Juk6NBIxUuHoPj0oH/2kotrl2nJtgAahs8fmwIv9BdAcHa1JSBbgooXwbwTT4iQT3pf8ZB2l9M8HSPkPbxV0qA0DMEFRlksTEoWq3RiDmxugC10+oK677jAgx9uXtsnJSALkqqCJyR6u7V2E7D9ED3FRkfVhIEaUeQYvmfM/Azg3IuerY1w9qRKDl/JvcrehajpHdE6/FY6xeAQ8S5HPRtt8FwHIECAuUWnowjRuL0IqeoxZJFqjrLfdbZdkdlwj/bKe7lmnXUNGhT3X65cr1RUSCUXtDBAci6AjOQqssBbNS73/9S7KYqK2BpBrkldSZ43uANCNySWtJySNXUGUE6Hm3vgmk6S3Op5VySGKxQ0o82MHmzYjCjwtPNA71AP1cvK1HNYzUrPZ3/mat7ejhpe2eNuVh3xzTK8gcXBVYEtTNNKzFwQp5psY8Zku+4gLY6XoEYUvmhEpADEvskiNphrRVtql9uwWdSY7mggdedkqoYi261zArbomj8eWMJpmiczvyUuk4UKBVwbn6yqcRsjldolMRdCg4lC1AnbM8kNYqyoQYtjfm42wmyQeuLwHzOYNp2oywIyWP1EkgaSF1EXdvrKuAl/lBP0DHjTmegcleIJQOD7RtANnZNKiL/Ar3v0lofQZMwXBEICEGL2RT3wByk9mnUCSXRzTetLmmyln/im5cDi6c0mJCwq3BdQX5gORy cZy1/Cj9 Pdd9jfMLwgioLljWUfgHeoqzlJu8mgDZjMOoIYXZQwjMLUDHDNaH2aW8xgN2xG3PjqhQP7aRT8MTUqoNXprKvNOBBYSeEiQn/gX13kxjX87r5Dk2VYI1s/dkvz7lIui9yEWJxMtkynPXxGSymMGZY/Mo1QfE+9CsdzDN0Ul7LtHGhJsHe2C3sT9YMLLzjFihk8PeTha07BZupz1UUpZ/G0Vhd05N/qNV4huBF2rJ0wDREzSQPPcs211GR8COHWgG4DE5te87E5cvru26Yyoi7XESD3f/LRSVNjQPTSY/QJE2ChL0= X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: From: Luis Chamberlain split_folio() and split_folio_to_list() assume order 0, to support minorder for non-anonymous folios, we must expand these to check the folio mapping order and use that. Set new_order to be at least minimum folio order if it is set in split_huge_page_to_list() so that we can maintain minimum folio order requirement in the page cache. Update the debugfs write files used for testing to ensure the order is respected as well. We simply enforce the min order when a file mapping is used. Signed-off-by: Luis Chamberlain Signed-off-by: Pankaj Raghav Reviewed-by: Hannes Reinecke --- There was a discussion about whether we need to consider truncation of folio to be considered a split failure or not [1]. The new code has retained the existing behaviour of returning a failure if the folio was truncated. I think we need to have a separate discussion whethere or not to consider it as a failure. include/linux/huge_mm.h | 14 ++++++++--- mm/huge_memory.c | 55 ++++++++++++++++++++++++++++++++++++++--- 2 files changed, 61 insertions(+), 8 deletions(-) diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index 212cca384d7e..70d80d17c3ff 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -90,6 +90,8 @@ extern struct kobj_attribute thpsize_shmem_enabled_attr; #define thp_vma_allowable_order(vma, vm_flags, tva_flags, order) \ (!!thp_vma_allowable_orders(vma, vm_flags, tva_flags, BIT(order))) +#define split_folio(f) split_folio_to_list(f, NULL) + #ifdef CONFIG_PGTABLE_HAS_HUGE_LEAVES #define HPAGE_PMD_SHIFT PMD_SHIFT #define HPAGE_PUD_SHIFT PUD_SHIFT @@ -320,9 +322,10 @@ unsigned long thp_get_unmapped_area_vmflags(struct file *filp, unsigned long add bool can_split_folio(struct folio *folio, int *pextra_pins); int split_huge_page_to_list_to_order(struct page *page, struct list_head *list, unsigned int new_order); +int split_folio_to_list(struct folio *folio, struct list_head *list); static inline int split_huge_page(struct page *page) { - return split_huge_page_to_list_to_order(page, NULL, 0); + return split_folio(page_folio(page)); } void deferred_split_folio(struct folio *folio); @@ -487,6 +490,12 @@ static inline int split_huge_page(struct page *page) { return 0; } + +static inline int split_folio_to_list(struct folio *folio, struct list_head *list) +{ + return 0; +} + static inline void deferred_split_folio(struct folio *folio) {} #define split_huge_pmd(__vma, __pmd, __address) \ do { } while (0) @@ -601,7 +610,4 @@ static inline int split_folio_to_order(struct folio *folio, int new_order) return split_folio_to_list_to_order(folio, NULL, new_order); } -#define split_folio_to_list(f, l) split_folio_to_list_to_order(f, l, 0) -#define split_folio(f) split_folio_to_order(f, 0) - #endif /* _LINUX_HUGE_MM_H */ diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 0fffaa58a47a..51fda5f9ac90 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3065,6 +3065,9 @@ bool can_split_folio(struct folio *folio, int *pextra_pins) * released, or if some unexpected race happened (e.g., anon VMA disappeared, * truncation). * + * Callers should ensure that the order respects the address space mapping + * min-order if one is set for non-anonymous folios. + * * Returns -EINVAL when trying to split to an order that is incompatible * with the folio. Splitting to order 0 is compatible with all folios. */ @@ -3146,6 +3149,7 @@ int split_huge_page_to_list_to_order(struct page *page, struct list_head *list, mapping = NULL; anon_vma_lock_write(anon_vma); } else { + unsigned int min_order; gfp_t gfp; mapping = folio->mapping; @@ -3156,6 +3160,14 @@ int split_huge_page_to_list_to_order(struct page *page, struct list_head *list, goto out; } + min_order = mapping_min_folio_order(folio->mapping); + if (new_order < min_order) { + VM_WARN_ONCE(1, "Cannot split mapped folio below min-order: %u", + min_order); + ret = -EINVAL; + goto out; + } + gfp = current_gfp_context(mapping_gfp_mask(mapping) & GFP_RECLAIM_MASK); @@ -3267,6 +3279,21 @@ int split_huge_page_to_list_to_order(struct page *page, struct list_head *list, return ret; } +int split_folio_to_list(struct folio *folio, struct list_head *list) +{ + unsigned int min_order = 0; + + if (!folio_test_anon(folio)) { + if (!folio->mapping) { + count_vm_event(THP_SPLIT_PAGE_FAILED); + return -EBUSY; + } + min_order = mapping_min_folio_order(folio->mapping); + } + + return split_huge_page_to_list_to_order(&folio->page, list, min_order); +} + void __folio_undo_large_rmappable(struct folio *folio) { struct deferred_split *ds_queue; @@ -3496,6 +3523,8 @@ static int split_huge_pages_pid(int pid, unsigned long vaddr_start, struct vm_area_struct *vma = vma_lookup(mm, addr); struct page *page; struct folio *folio; + struct address_space *mapping; + unsigned int target_order = new_order; if (!vma) break; @@ -3516,7 +3545,13 @@ static int split_huge_pages_pid(int pid, unsigned long vaddr_start, if (!is_transparent_hugepage(folio)) goto next; - if (new_order >= folio_order(folio)) + if (!folio_test_anon(folio)) { + mapping = folio->mapping; + target_order = max(new_order, + mapping_min_folio_order(mapping)); + } + + if (target_order >= folio_order(folio)) goto next; total++; @@ -3532,9 +3567,13 @@ static int split_huge_pages_pid(int pid, unsigned long vaddr_start, if (!folio_trylock(folio)) goto next; - if (!split_folio_to_order(folio, new_order)) + if (!folio_test_anon(folio) && folio->mapping != mapping) + goto unlock; + + if (!split_folio_to_order(folio, target_order)) split++; +unlock: folio_unlock(folio); next: folio_put(folio); @@ -3559,6 +3598,7 @@ static int split_huge_pages_in_file(const char *file_path, pgoff_t off_start, pgoff_t index; int nr_pages = 1; unsigned long total = 0, split = 0; + unsigned int min_order; file = getname_kernel(file_path); if (IS_ERR(file)) @@ -3572,9 +3612,11 @@ static int split_huge_pages_in_file(const char *file_path, pgoff_t off_start, file_path, off_start, off_end); mapping = candidate->f_mapping; + min_order = mapping_min_folio_order(mapping); for (index = off_start; index < off_end; index += nr_pages) { struct folio *folio = filemap_get_folio(mapping, index); + unsigned int target_order = new_order; nr_pages = 1; if (IS_ERR(folio)) @@ -3583,18 +3625,23 @@ static int split_huge_pages_in_file(const char *file_path, pgoff_t off_start, if (!folio_test_large(folio)) goto next; + target_order = max(new_order, min_order); total++; nr_pages = folio_nr_pages(folio); - if (new_order >= folio_order(folio)) + if (target_order >= folio_order(folio)) goto next; if (!folio_trylock(folio)) goto next; - if (!split_folio_to_order(folio, new_order)) + if (folio->mapping != mapping) + goto unlock; + + if (!split_folio_to_order(folio, target_order)) split++; +unlock: folio_unlock(folio); next: folio_put(folio);