[v18,32/32] mm: Split release_pages work into 3 passes

Message ID	1598273705-69124-33-git-send-email-alex.shi@linux.alibaba.com (mailing list archive)
State	New, archived
Headers	show Return-Path: <SRS0=JPLW=CC=kvack.org=owner-linux-mm@kernel.org> DMARC-Filter: OpenDMARC Filter v1.3.2 mail.kernel.org DFB0C20838 From: Alex Shi <alex.shi@linux.alibaba.com> To: akpm@linux-foundation.org, mgorman@techsingularity.net, tj@kernel.org, hughd@google.com, khlebnikov@yandex-team.ru, daniel.m.jordan@oracle.com, willy@infradead.org, hannes@cmpxchg.org, lkp@intel.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, shakeelb@google.com, iamjoonsoo.kim@lge.com, richard.weiyang@gmail.com, kirill@shutemov.name, alexander.duyck@gmail.com, rong.a.chen@intel.com, mhocko@suse.com, vdavydov.dev@gmail.com, shy828301@gmail.com Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com> Subject: [PATCH v18 32/32] mm: Split release_pages work into 3 passes Date: Mon, 24 Aug 2020 20:55:05 +0800 Message-Id: <1598273705-69124-33-git-send-email-alex.shi@linux.alibaba.com> In-Reply-To: <1598273705-69124-1-git-send-email-alex.shi@linux.alibaba.com> References: <1598273705-69124-1-git-send-email-alex.shi@linux.alibaba.com> Sender: owner-linux-mm@kvack.org Precedence: bulk
Series	per memcg lru_lock \| expand [v18,00/32] per memcg lru_lock [v18,01/32] mm/memcg: warning on !memcg after readahead page charged [v18,02/32] mm/memcg: bail out early from swap accounting when memcg is disabled [v18,03/32] mm/thp: move lru_add_page_tail func to huge_memory.c [v18,04/32] mm/thp: clean up lru_add_page_tail [v18,05/32] mm/thp: remove code path which never got into [v18,06/32] mm/thp: narrow lru locking [v18,07/32] mm/swap.c: stop deactivate_file_page if page not on lru [v18,08/32] mm/vmscan: remove unnecessary lruvec adding [v18,09/32] mm/page_idle: no unlikely double check for idle page counting [v18,10/32] mm/compaction: rename compact_deferred as compact_should_defer [v18,11/32] mm/memcg: add debug checking in lock_page_memcg [v18,12/32] mm/memcg: optimize mem_cgroup_page_lruvec [v18,13/32] mm/swap.c: fold vm event PGROTATED into pagevec_move_tail_fn [v18,14/32] mm/lru: move lru_lock holding in func lru_note_cost_page [v18,15/32] mm/lru: move lock into lru_note_cost [v18,16/32] mm/lru: introduce TestClearPageLRU [v18,17/32] mm/compaction: do page isolation first in compaction [v18,18/32] mm/thp: add tail pages into lru anyway in split_huge_page() [v18,19/32] mm/swap.c: serialize memcg changes in pagevec_lru_move_fn [v18,20/32] mm/lru: replace pgdat lru_lock with lruvec lock [v18,21/32] mm/lru: introduce the relock_page_lruvec function [v18,22/32] mm/vmscan: use relock for move_pages_to_lru [v18,23/32] mm/lru: revise the comments of lru_lock [v18,24/32] mm/pgdat: remove pgdat lru_lock [v18,25/32] mm/mlock: remove lru_lock on TestClearPageMlocked in munlock_vma_page [v18,26/32] mm/mlock: remove __munlock_isolate_lru_page [v18,27/32] mm/swap.c: optimizing __pagevec_lru_add lru_lock [v18,28/32] mm/compaction: Drop locked from isolate_migratepages_block [v18,29/32] mm: Identify compound pages sooner in isolate_migratepages_block [v18,30/32] mm: Drop use of test_and_set_skip in favor of just setting skip [v18,31/32] mm: Add explicit page decrement in exception path for isolate_lru_pages [v18,32/32] mm: Split release_pages work into 3 passes

Message ID

1598273705-69124-33-git-send-email-alex.shi@linux.alibaba.com (mailing list archive)

State

New, archived

Headers

DMARC-Filter: OpenDMARC Filter v1.3.2 mail.kernel.org DFB0C20838
From: Alex Shi <alex.shi@linux.alibaba.com>
To: akpm@linux-foundation.org,
	mgorman@techsingularity.net,
	tj@kernel.org,
	hughd@google.com,
	khlebnikov@yandex-team.ru,
	daniel.m.jordan@oracle.com,
	willy@infradead.org,
	hannes@cmpxchg.org,
	lkp@intel.com,
	linux-mm@kvack.org,
	linux-kernel@vger.kernel.org,
	cgroups@vger.kernel.org,
	shakeelb@google.com,
	iamjoonsoo.kim@lge.com,
	richard.weiyang@gmail.com,
	kirill@shutemov.name,
	alexander.duyck@gmail.com,
	rong.a.chen@intel.com,
	mhocko@suse.com,
	vdavydov.dev@gmail.com,
	shy828301@gmail.com
Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Subject: [PATCH v18 32/32] mm: Split release_pages work into 3 passes
Date: Mon, 24 Aug 2020 20:55:05 +0800
Message-Id: <1598273705-69124-33-git-send-email-alex.shi@linux.alibaba.com>
In-Reply-To: <1598273705-69124-1-git-send-email-alex.shi@linux.alibaba.com>
References: <1598273705-69124-1-git-send-email-alex.shi@linux.alibaba.com>
Sender: owner-linux-mm@kvack.org
Precedence: bulk

Series

per memcg lru_lock | expand

Commit Message

Alex Shi Aug. 24, 2020, 12:55 p.m. UTC

From: Alexander Duyck <alexander.h.duyck@linux.intel.com>

The release_pages function has a number of paths that end up with the
LRU lock having to be released and reacquired. Such an example would be the
freeing of THP pages as it requires releasing the LRU lock so that it can
be potentially reacquired by __put_compound_page.

In order to avoid that we can split the work into 3 passes, the first
without the LRU lock to go through and sort out those pages that are not in
the LRU so they can be freed immediately from those that can't. The second
pass will then go through removing those pages from the LRU in batches as
large as a pagevec can hold before freeing the LRU lock. Once the pages have
been removed from the LRU we can then proceed to free the remaining pages
without needing to worry about if they are in the LRU any further.

The general idea is to avoid bouncing the LRU lock between pages and to
hopefully aggregate the lock for up to the full page vector worth of pages.

Signed-off-by: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: linux-kernel@vger.kernel.org
Cc: linux-mm@kvack.org
---
 mm/swap.c | 109 ++++++++++++++++++++++++++++++++++++++------------------------
 1 file changed, 67 insertions(+), 42 deletions(-)

diff --git a/mm/swap.c b/mm/swap.c
index fe53449fa1b8..b405f81b2c60 100644
--- a/mm/swap.c
+++ b/mm/swap.c
@@ -795,6 +795,54 @@  void lru_add_drain_all(void)
 }
 #endif
 
+static void __release_page(struct page *page, struct list_head *pages_to_free)
+{
+	if (PageCompound(page)) {
+		__put_compound_page(page);
+	} else {
+		/* Clear Active bit in case of parallel mark_page_accessed */
+		__ClearPageActive(page);
+		__ClearPageWaiters(page);
+
+		list_add(&page->lru, pages_to_free);
+	}
+}
+
+static void __release_lru_pages(struct pagevec *pvec,
+				struct list_head *pages_to_free)
+{
+	struct lruvec *lruvec = NULL;
+	unsigned long flags = 0;
+	int i;
+
+	/*
+	 * The pagevec at this point should contain a set of pages with
+	 * their reference count at 0 and the LRU flag set. We will now
+	 * need to pull the pages from their LRU lists.
+	 *
+	 * We walk the list backwards here since that way we are starting at
+	 * the pages that should be warmest in the cache.
+	 */
+	for (i = pagevec_count(pvec); i--;) {
+		struct page *page = pvec->pages[i];
+
+		lruvec = relock_page_lruvec_irqsave(page, lruvec, &flags);
+		VM_BUG_ON_PAGE(!PageLRU(page), page);
+		__ClearPageLRU(page);
+		del_page_from_lru_list(page, lruvec, page_off_lru(page));
+	}
+
+	unlock_page_lruvec_irqrestore(lruvec, flags);
+
+	/*
+	 * A batch of pages are no longer on the LRU list. Go through and
+	 * start the final process of returning the deferred pages to their
+	 * appropriate freelists.
+	 */
+	for (i = pagevec_count(pvec); i--;)
+		__release_page(pvec->pages[i], pages_to_free);
+}
+
 /**
  * release_pages - batched put_page()
  * @pages: array of pages to release
@@ -806,32 +854,24 @@  void lru_add_drain_all(void)
 void release_pages(struct page **pages, int nr)
 {
 	int i;
+	struct pagevec pvec;
 	LIST_HEAD(pages_to_free);
-	struct lruvec *lruvec = NULL;
-	unsigned long flags;
-	unsigned int lock_batch;
 
+	pagevec_init(&pvec);
+
+	/*
+	 * We need to first walk through the list cleaning up the low hanging
+	 * fruit and clearing those pages that either cannot be freed or that
+	 * are non-LRU. We will store the LRU pages in a pagevec so that we
+	 * can get to them in the next pass.
+	 */
 	for (i = 0; i < nr; i++) {
 		struct page *page = pages[i];
 
-		/*
-		 * Make sure the IRQ-safe lock-holding time does not get
-		 * excessive with a continuous string of pages from the
-		 * same lruvec. The lock is held only if lruvec != NULL.
-		 */
-		if (lruvec && ++lock_batch == SWAP_CLUSTER_MAX) {
-			unlock_page_lruvec_irqrestore(lruvec, flags);
-			lruvec = NULL;
-		}
-
 		if (is_huge_zero_page(page))
 			continue;
 
 		if (is_zone_device_page(page)) {
-			if (lruvec) {
-				unlock_page_lruvec_irqrestore(lruvec, flags);
-				lruvec = NULL;
-			}
 			/*
 			 * ZONE_DEVICE pages that return 'false' from
 			 * put_devmap_managed_page() do not require special
@@ -848,36 +888,21 @@  void release_pages(struct page **pages, int nr)
 		if (!put_page_testzero(page))
 			continue;
 
-		if (PageCompound(page)) {
-			if (lruvec) {
-				unlock_page_lruvec_irqrestore(lruvec, flags);
-				lruvec = NULL;
-			}
-			__put_compound_page(page);
+		if (!PageLRU(page)) {
+			__release_page(page, &pages_to_free);
 			continue;
 		}
 
-		if (PageLRU(page)) {
-			struct lruvec *prev_lruvec = lruvec;
-
-			lruvec = relock_page_lruvec_irqsave(page, lruvec,
-									&flags);
-			if (prev_lruvec != lruvec)
-				lock_batch = 0;
-
-			VM_BUG_ON_PAGE(!PageLRU(page), page);
-			__ClearPageLRU(page);
-			del_page_from_lru_list(page, lruvec, page_off_lru(page));
+		/* record page so we can get it in the next pass */
+		if (!pagevec_add(&pvec, page)) {
+			__release_lru_pages(&pvec, &pages_to_free);
+			pagevec_reinit(&pvec);
 		}
-
-		/* Clear Active bit in case of parallel mark_page_accessed */
-		__ClearPageActive(page);
-		__ClearPageWaiters(page);
-
-		list_add(&page->lru, &pages_to_free);
 	}
-	if (lruvec)
-		unlock_page_lruvec_irqrestore(lruvec, flags);
+
+	/* flush any remaining LRU pages that need to be processed */
+	if (pagevec_count(&pvec))
+		__release_lru_pages(&pvec, &pages_to_free);
 
 	mem_cgroup_uncharge_list(&pages_to_free);
 	free_unref_page_list(&pages_to_free);

[v18,32/32] mm: Split release_pages work into 3 passes

Commit Message

Patch