From patchwork Thu Sep 21 08:10:57 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Yosry Ahmed X-Patchwork-Id: 13393762 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 351CDE706E5 for ; Thu, 21 Sep 2023 08:11:17 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id D00A46B01FB; Thu, 21 Sep 2023 04:11:13 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id CB38C6B01FE; Thu, 21 Sep 2023 04:11:13 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id ADBCF6B01FF; Thu, 21 Sep 2023 04:11:13 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) by kanga.kvack.org (Postfix) with ESMTP id 958D06B01FB for ; Thu, 21 Sep 2023 04:11:13 -0400 (EDT) Received: from smtpin15.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay02.hostedemail.com (Postfix) with ESMTP id 650F412100E for ; Thu, 21 Sep 2023 08:11:13 +0000 (UTC) X-FDA: 81259884426.15.E037921 Received: from mail-yb1-f201.google.com (mail-yb1-f201.google.com [209.85.219.201]) by imf18.hostedemail.com (Postfix) with ESMTP id 91DBB1C0009 for ; Thu, 21 Sep 2023 08:11:11 +0000 (UTC) Authentication-Results: imf18.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=sM1NhMT9; spf=pass (imf18.hostedemail.com: domain of 3nvoLZQoKCPErhlkrTafXWZhhZeX.Vhfebgnq-ffdoTVd.hkZ@flex--yosryahmed.bounces.google.com designates 209.85.219.201 as permitted sender) smtp.mailfrom=3nvoLZQoKCPErhlkrTafXWZhhZeX.Vhfebgnq-ffdoTVd.hkZ@flex--yosryahmed.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1695283871; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=3167ObCwuPGy0ykS2mGpCnZZz/lnDX5wo2e/FWfZcMc=; b=f5c63QvYy3cHwhuFH/SBCEelGx5DUimx6LAWkz9uNSUnU6fRkE9e4RCRP2HOWmQaBX6c42 q9iGzdYmNoaHW/ld2HqvKmR0AW+O/r2J8dmMY2lPYF8146obRusMiGa9lKYgpMBFPF+Z/g +Vcpmbk/G7I9lcy0VSeP3sdWLB6qKi8= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1695283871; a=rsa-sha256; cv=none; b=CVJV1Vj8PXngukeZR5NNggNV9F0YaxZFqjSZ93EO5B3Wo/w2yTVQwNE3uwK2aQVI1yoc7w p2BKrXnh7pvgxdXNsc9beo+ZR9RCfgaw9Ybo4VHR3Y/lCEv7ivzw/SLUbZ5pYJT13gWb76 YaORwiUH1gWuhOX4M4bXrXyjuxcBnJo= ARC-Authentication-Results: i=1; imf18.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=sM1NhMT9; spf=pass (imf18.hostedemail.com: domain of 3nvoLZQoKCPErhlkrTafXWZhhZeX.Vhfebgnq-ffdoTVd.hkZ@flex--yosryahmed.bounces.google.com designates 209.85.219.201 as permitted sender) smtp.mailfrom=3nvoLZQoKCPErhlkrTafXWZhhZeX.Vhfebgnq-ffdoTVd.hkZ@flex--yosryahmed.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com Received: by mail-yb1-f201.google.com with SMTP id 3f1490d57ef6-d81d85aae7cso3443001276.0 for ; Thu, 21 Sep 2023 01:11:11 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1695283870; x=1695888670; darn=kvack.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=3167ObCwuPGy0ykS2mGpCnZZz/lnDX5wo2e/FWfZcMc=; b=sM1NhMT9Q1QiU5fVpdXXyvMGhdOPvs+ZLFZQiAgs1gP++oSYXq+J0HivrbUqHpc9zg OG0xQvChS0M9fWBZmLCCeX13HBm/C8E4G2H9t1rzTgzkeO0r9WpUuMIHy7OovvcSWg4+ 7MKXYQwJe7IT9mXr8p8gntJUHC2WZlpyiWC7+A6V8pKGYygJHA0adLT9r7hpROdY4FF1 bzPNgHyzVCiIoNmJHwNxaZzAgneXaLIY0lxaVAngNUUWmvIt7rQa/4+PQc6GCG0VyAiL t6Mt7SNF80XLXZTgDhdAypvTFqyvbg3K6sPQ5qcPTdSqO3OwelvTi5DNx5Gqn+voSEgM L2NA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1695283870; x=1695888670; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=3167ObCwuPGy0ykS2mGpCnZZz/lnDX5wo2e/FWfZcMc=; b=DZCHOegFA84MmKgp2ZLKLMUKNG7D9UBnELw8bOPbNrQ0bME1mUXNZvF2AWHvim2TWZ bUko9pIdcOU9xuIdcvsveKH5fE/dAsmN+9hoa0S6XDnSc7F5KlzhpWVx5h1BRiEtVwkB rhUy+Dh+zi5QNSYJlcYgfrVHI9rdGlFIZzmKTnDU6gQbBqPURedSpjWQ4FnLBY1AKj8i QC54U5011GuC+TK+IEQrBCjRjCv3UV4ysG66M+GIrjzxlJj60i+FuXv+qOh0sOVlwXnt 9dA0KvS+YUenUN+kHY8qbnHCRkfaXbI2GnhnVXc9XofkWYvJYr4shNIq/Wta+T1idVgm PZFQ== X-Gm-Message-State: AOJu0YxQmn4GD8hgwSDbA/YLx0NyMenJdoZwGQXYr6GQTFZhsPayd1wp 1PDevt3xmcYW6CwNbd12Bo4VorVgcDGIAjK8 X-Google-Smtp-Source: AGHT+IEJ4fcoyEGYjUXnEO4k8S62qHUkVPnBwaKgs95Sypxteexrw0RSMgrAyDrgXkVHLcy4yrd/Nrxvhz1akcJT X-Received: from yosry.c.googlers.com ([fda3:e722:ac3:cc00:20:ed76:c0a8:29b4]) (user=yosryahmed job=sendgmr) by 2002:a25:e0c5:0:b0:d78:28d0:15bc with SMTP id x188-20020a25e0c5000000b00d7828d015bcmr99461ybg.4.1695283870798; Thu, 21 Sep 2023 01:11:10 -0700 (PDT) Date: Thu, 21 Sep 2023 08:10:57 +0000 In-Reply-To: <20230921081057.3440885-1-yosryahmed@google.com> Mime-Version: 1.0 References: <20230921081057.3440885-1-yosryahmed@google.com> X-Mailer: git-send-email 2.42.0.459.ge4e396fd5e-goog Message-ID: <20230921081057.3440885-6-yosryahmed@google.com> Subject: [PATCH 5/5] mm: memcg: restore subtree stats flushing From: Yosry Ahmed To: Andrew Morton Cc: Johannes Weiner , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Ivan Babrou , Tejun Heo , " =?utf-8?q?Michal_Koutn=C3=BD?= " , Waiman Long , kernel-team@cloudflare.com, Wei Xu , Greg Thelen , linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, Yosry Ahmed X-Rspamd-Queue-Id: 91DBB1C0009 X-Rspam-User: X-Stat-Signature: 49zmnmsbqf4bw43ymgza9cjfcwre4bp5 X-Rspamd-Server: rspam03 X-HE-Tag: 1695283871-416646 X-HE-Meta: U2FsdGVkX1+dGH64S46fPV44SLo9oRiVFxSjoKcHxu6zn4PGw8tu1Cuaxnov204GZBUhBr+EQ4tm8aICNEpk4Bv7ts3yBCqLgzy8ROtZayuku/0AeFGirfCQhlTq07XQ3u9eZWG0ZeCadZk6fTEjyuXgHHY9c+Faw+GknpL3qrKFEH4tDPTmUbfifrSMtzrS3XOet55U3ZzGKCFSWnQJL8FYg7QDhtJRuqJRdCryvUr5/cFlXhkQK8ujouMn1XzbJQfbSIF3MgziG4dqKRy8YrhWG+vk3IIrqbVmxTB2DF7ojnJOre3HGcqFEGfchT7euZUIzuJmGDgFCY5GM7HRu9A2FYUKsl2/DqyQeLbSKJsoOPYCuhe6mofhoYbLjx/dSlMT+K2u8Y3SnZDWKxuVHZ+0rlys7u2ynuDtofQR6i5+oUkAyUdYGG6E4OiKCToGZ13235iyKX2PKQYZtG863amNUBARw6oz5Nf9xxQcaiDTGe75/wKd8qIlwuWY5+i5ZT+diEjA4H6M2I3k8dWBPTBeO9SNd4I5Tvgijr1QhfNPDjBzgykxWPCAlxwh8GfKE5lT/sJ+pn1B9oZssqG65T0Uwe4L65Pmfr3GWporAx9WNFRDqDGRxEIHqa4bgZfIR6kXypzQVoOI1RQRMKagxUFaFDMz+AMrzS2aygR7eIm3ktuCQP1xk0D4GzAbCj9KauMytIfxZZ6E31dTbrHIKtJIMGZZ745kcefftDRDDw9i2xcX5Z9E5zRUkZMYYfmxM7k8CPpC0A1C4CXCY3vSefanOyW24/z6GrtPNUxikDKNpZeh56eM0hjKUMXstywj9hWfxJgodr9nokHedoZXWaO0vaph/0CesSPU9/6s9o4SKnT27NOzupsHd6npwVOFJx/7ZXT8o4uQtGZ0D+IJ22WI4NTRl+dHxEWGYzfJznno8uyEqPvsHMU3/FNFmnnED+kGlyXFish8KZ6n5VZ 5A5tMxch psqWkR5tre727+nGvctAX4xuDIMtrGzrHhwyfwQoZdOC2+GId70ExbT+zmCvuTzbGFcMdDEVjOdQz/6XpsxKfhqDvtQkHR58SCFQ3Pfe5BGgXKRSHcFtfZd+7i0AYSzysS1QMeTGcqTJZTSVCc5lKTmOFp2oB/cb8AwNIRfsfK1tapZ2ekAslrDhLqeJleeQQaDMLnhUqqijitbUNnILgCeeYxcAAr/Su6EMA6QTo1Hb6/ylsavWGVFavDwKwXno8lTDkvRdz7ql5fGlNZlm/3LuWtmTLG7D8XTlq+V+HLoXLGGUgiwgpHhW6n1eiwDUmEpOIfg6NsBZnQq3Wrr7rvV8tDF+gQf63k9LDYN4VkBr3s28hX95KBZvVQe+AvHor0UCSP2HzHMYEhlNrtBCwQG5Nf8YbaxTDFVv3k4AdJpOwX65OMgMnTmmuDxI8MrfiFNKZPKp8qBrXpKf7gajvQAws0rCPsVhb2RY8LkjPSiAf6B9FAmnl2gRSu0LGoEK+9uy5gc+CS2rl32c3e7p5Vt600+Ph0Qm1IR7e3eSuqc+TdIfPP862hp9L6P8V5jgmZFg3dZ+nhXmvGYhkXQD7OO/JiI9BCcYGCE14NsRlpmY3zXJyp1qSP7H1DB8KkOPS0EYosPNcpZrELmHuOb9VkMYUK5DJcUttAkmAZgq1oyThEOWSrJlDlSmpfK/YT3K2OTFLaibNJUmwpvmlzh+U6zksnJKdlnNqXQ0E X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: Stats flushing for memcg currently follows the following rules: - Always flush the entire memcg hierarchy (i.e. flush the root). - Only one flusher is allowed at a time. If someone else tries to flush concurrently, they skip and return immediately. - A periodic flusher flushes all the stats every 2 seconds. The reason this approach is followed is because all flushes are serialized by a global rstat spinlock. On the memcg side, flushing is invoked from userspace reads as well as in-kernel flushers (e.g. reclaim, refault, etc). This approach aims to avoid serializing all flushers on the global lock, which can cause a significant performance hit under high concurrency. This approach has the following problems: - Occasionally a userspace read of the stats of a non-root cgroup will be too expensive as it has to flush the entire hierarchy [1]. - Sometimes the stats accuracy are compromised if there is an ongoing flush, and we skip and return before the subtree of interest is actually flushed, yielding stale stats (by up to 2s due to periodic flushing). This is more visible when reading stats from userspace, but can also affect in-kernel flushers. The latter problem is particulary a concern when userspace reads stats after an event occurs, but gets stats from before the event. Examples: - When memory usage / pressure spikes, a userspace OOM handler may look at the stats of different memcgs to select a victim based on various heuristics (e.g. how much private memory will be freed by killing this). Reading stale stats from before the usage spike in this case may cause a wrongful OOM kill. - A proactive reclaimer may read the stats after writing to memory.reclaim to measure the success of the reclaim operation. Stale stats from before reclaim may give a false negative. - Reading the stats of a parent and a child memcg may be inconsistent (child larger than parent), if the flush doesn't happen when the parent is read, but happens when the child is read. As for in-kernel flushers, they will occasionally get stale stats. No regressions are currently known from this, but if there are regressions, they would be very difficult to debug and link to the source of the problem. This patch aims to fix these problems by restoring subtree flushing, and removing the unified/coalesced flushing logic that skips flushing if there is an ongoing flush. This change would introduce a significant regression with global stats flushing thresholds. With per-memcg stats flushing thresholds, this seems to perform really well. The thresholds protect the underlying lock from unnecessary contention. Add a mutex to protect the underlying rstat lock from excessive memcg flushing. The thresholds are re-checked after the mutex is grabbed to make sure that a concurrent flush did not already get the subtree we are trying to flush. A call to cgroup_rstat_flush() is not cheap, even if there are no pending updates. This patch was tested in two ways to ensure the latency of flushing is up to bar, on a machine with 384 cpus: - A synthetic test with 5000 concurrent workers in 500 cgroups doing allocations and reclaim, as well as 1000 readers for memory.stat (variation of [2]). No regressions were noticed in the total runtime. Note that significant regressions in this test are observed with global stats thresholds, but not with per-memcg thresholds. - A synthetic stress test for concurrently reading memcg stats while memory allocation/freeing workers are running in the background, provided by Wei Xu [3]. With 250k threads reading the stats every 100ms in 50k cgroups, 99.9% of reads take <= 50us. Less than 0.01% of reads take more than 1ms, and no reads take more than 100ms. [1] https://lore.kernel.org/lkml/CABWYdi0c6__rh-K7dcM_pkf9BJdTRtAU08M43KO9ME4-dsgfoQ@mail.gmail.com/ [2] https://lore.kernel.org/lkml/CAJD7tka13M-zVZTyQJYL1iUAYvuQ1fcHbCjcOBZcz6POYTV-4g@mail.gmail.com/ [3] https://lore.kernel.org/lkml/CAAPL-u9D2b=iF5Lf_cRnKxUfkiEe0AMDTu6yhrUAzX0b6a6rDg@mail.gmail.com/ Signed-off-by: Yosry Ahmed --- include/linux/memcontrol.h | 8 ++--- mm/memcontrol.c | 73 +++++++++++++++++++++++--------------- mm/vmscan.c | 2 +- mm/workingset.c | 10 ++++-- 4 files changed, 56 insertions(+), 37 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 45d0c10e86cc..1b61a2707307 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -1030,8 +1030,8 @@ static inline unsigned long lruvec_page_state_local(struct lruvec *lruvec, return x; } -void mem_cgroup_flush_stats(void); -void mem_cgroup_flush_stats_ratelimited(void); +void mem_cgroup_flush_stats(struct mem_cgroup *memcg); +void mem_cgroup_flush_stats_ratelimited(struct mem_cgroup *memcg); void __mod_memcg_lruvec_state(struct lruvec *lruvec, enum node_stat_item idx, int val); @@ -1520,11 +1520,11 @@ static inline unsigned long lruvec_page_state_local(struct lruvec *lruvec, return node_page_state(lruvec_pgdat(lruvec), idx); } -static inline void mem_cgroup_flush_stats(void) +static inline void mem_cgroup_flush_stats(struct mem_cgroup *memcg) { } -static inline void mem_cgroup_flush_stats_ratelimited(void) +static inline void mem_cgroup_flush_stats_ratelimited(struct mem_cgroup *memcg) { } diff --git a/mm/memcontrol.c b/mm/memcontrol.c index c273c65bb642..99cfba81684f 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -666,7 +666,6 @@ struct memcg_vmstats { */ static void flush_memcg_stats_dwork(struct work_struct *w); static DECLARE_DEFERRABLE_WORK(stats_flush_dwork, flush_memcg_stats_dwork); -static atomic_t stats_flush_ongoing = ATOMIC_INIT(0); static u64 flush_last_time; #define FLUSH_TIME (2UL*HZ) @@ -727,35 +726,45 @@ static inline void memcg_rstat_updated(struct mem_cgroup *memcg, int val) } } -static void do_flush_stats(void) +static void do_flush_stats(struct mem_cgroup *memcg) { - /* - * We always flush the entire tree, so concurrent flushers can just - * skip. This avoids a thundering herd problem on the rstat global lock - * from memcg flushers (e.g. reclaim, refault, etc). - */ - if (atomic_read(&stats_flush_ongoing) || - atomic_xchg(&stats_flush_ongoing, 1)) - return; - - WRITE_ONCE(flush_last_time, jiffies_64); - - cgroup_rstat_flush(root_mem_cgroup->css.cgroup); + if (mem_cgroup_is_root(memcg)) + WRITE_ONCE(flush_last_time, jiffies_64); - atomic_set(&stats_flush_ongoing, 0); + cgroup_rstat_flush(memcg->css.cgroup); } -void mem_cgroup_flush_stats(void) +/* + * mem_cgroup_flush_stats - flush the stats of a memory cgroup subtree + * @memcg: root of the subtree to flush + * + * Flushing is serialized by the underlying global rstat lock. There is also a + * minimum amount of work to be done even if there are no stat updates to flush. + * Hence, we only flush the stats if the updates delta exceeds a threshold. This + * avoids unnecessary work and contention on the underlying lock. + */ +void mem_cgroup_flush_stats(struct mem_cgroup *memcg) { - if (memcg_should_flush_stats(root_mem_cgroup)) - do_flush_stats(); + static DEFINE_MUTEX(memcg_stats_flush_mutex); + + if (!memcg) + memcg = root_mem_cgroup; + + if (!memcg_should_flush_stats(memcg)) + return; + + mutex_lock(&memcg_stats_flush_mutex); + /* An overlapping flush may have occurred, check again after locking */ + if (memcg_should_flush_stats(memcg)) + do_flush_stats(memcg); + mutex_unlock(&memcg_stats_flush_mutex); } -void mem_cgroup_flush_stats_ratelimited(void) +void mem_cgroup_flush_stats_ratelimited(struct mem_cgroup *memcg) { /* Only flush if the periodic flusher is one full cycle late */ if (time_after64(jiffies_64, READ_ONCE(flush_last_time) + 2*FLUSH_TIME)) - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(memcg); } static void flush_memcg_stats_dwork(struct work_struct *w) @@ -764,7 +773,7 @@ static void flush_memcg_stats_dwork(struct work_struct *w) * Deliberately ignore memcg_should_flush_stats() here so that flushing * in latency-sensitive paths is as cheap as possible. */ - do_flush_stats(); + do_flush_stats(root_mem_cgroup); queue_delayed_work(system_unbound_wq, &stats_flush_dwork, FLUSH_TIME); } @@ -1593,7 +1602,7 @@ static void memcg_stat_format(struct mem_cgroup *memcg, struct seq_buf *s) * * Current memory state: */ - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(memcg); for (i = 0; i < ARRAY_SIZE(memory_stats); i++) { u64 size; @@ -4035,7 +4044,7 @@ static int memcg_numa_stat_show(struct seq_file *m, void *v) int nid; struct mem_cgroup *memcg = mem_cgroup_from_seq(m); - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(memcg); for (stat = stats; stat < stats + ARRAY_SIZE(stats); stat++) { seq_printf(m, "%s=%lu", stat->name, @@ -4116,7 +4125,7 @@ static void memcg1_stat_format(struct mem_cgroup *memcg, struct seq_buf *s) BUILD_BUG_ON(ARRAY_SIZE(memcg1_stat_names) != ARRAY_SIZE(memcg1_stats)); - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(memcg); for (i = 0; i < ARRAY_SIZE(memcg1_stats); i++) { unsigned long nr; @@ -4613,7 +4622,7 @@ void mem_cgroup_wb_stats(struct bdi_writeback *wb, unsigned long *pfilepages, struct mem_cgroup *memcg = mem_cgroup_from_css(wb->memcg_css); struct mem_cgroup *parent; - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(memcg); *pdirty = memcg_page_state(memcg, NR_FILE_DIRTY); *pwriteback = memcg_page_state(memcg, NR_WRITEBACK); @@ -6640,7 +6649,7 @@ static int memory_numa_stat_show(struct seq_file *m, void *v) int i; struct mem_cgroup *memcg = mem_cgroup_from_seq(m); - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(memcg); for (i = 0; i < ARRAY_SIZE(memory_stats); i++) { int nid; @@ -7806,7 +7815,11 @@ bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) break; } - cgroup_rstat_flush(memcg->css.cgroup); + /* + * mem_cgroup_flush_stats() ignores small changes. Use + * do_flush_stats() directly to get accurate stats for charging. + */ + do_flush_stats(memcg); pages = memcg_page_state(memcg, MEMCG_ZSWAP_B) / PAGE_SIZE; if (pages < max) continue; @@ -7871,8 +7884,10 @@ void obj_cgroup_uncharge_zswap(struct obj_cgroup *objcg, size_t size) static u64 zswap_current_read(struct cgroup_subsys_state *css, struct cftype *cft) { - cgroup_rstat_flush(css->cgroup); - return memcg_page_state(mem_cgroup_from_css(css), MEMCG_ZSWAP_B); + struct mem_cgroup *memcg = mem_cgroup_from_css(css); + + mem_cgroup_flush_stats(memcg); + return memcg_page_state(memcg, MEMCG_ZSWAP_B); } static int zswap_max_show(struct seq_file *m, void *v) diff --git a/mm/vmscan.c b/mm/vmscan.c index a4e44f1c97c1..60bead17b1f7 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -2246,7 +2246,7 @@ static void prepare_scan_control(pg_data_t *pgdat, struct scan_control *sc) * Flush the memory cgroup stats, so that we read accurate per-memcg * lruvec stats for heuristics. */ - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(sc->target_mem_cgroup); /* * Determine the scan balance between anon and file LRUs. diff --git a/mm/workingset.c b/mm/workingset.c index 79b338996088..141defbe3da2 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -461,8 +461,12 @@ bool workingset_test_recent(void *shadow, bool file, bool *workingset) css_get(&eviction_memcg->css); rcu_read_unlock(); - /* Flush stats (and potentially sleep) outside the RCU read section */ - mem_cgroup_flush_stats_ratelimited(); + /* + * Flush stats (and potentially sleep) outside the RCU read section. + * XXX: With per-memcg flushing and thresholding, is ratelimiting + * still needed here? + */ + mem_cgroup_flush_stats_ratelimited(eviction_memcg); eviction_lruvec = mem_cgroup_lruvec(eviction_memcg, pgdat); refault = atomic_long_read(&eviction_lruvec->nonresident_age); @@ -673,7 +677,7 @@ static unsigned long count_shadow_nodes(struct shrinker *shrinker, struct lruvec *lruvec; int i; - mem_cgroup_flush_stats(); + mem_cgroup_flush_stats(sc->memcg); lruvec = mem_cgroup_lruvec(sc->memcg, NODE_DATA(sc->nid)); for (pages = 0, i = 0; i < NR_LRU_LISTS; i++) pages += lruvec_page_state_local(lruvec,