[v3,1/7] rcu: Reduce synchronize_rcu() latency

A call to a synchronize_rcu() can be optimized from a latency
point of view. Workloads which depend on this can benefit of it.

The delay of wakeme_after_rcu() callback, which unblocks a waiter,
depends on several factors:

- how fast a process of offloading is started. Combination of:
    - !CONFIG_RCU_NOCB_CPU/CONFIG_RCU_NOCB_CPU;
    - !CONFIG_RCU_LAZY/CONFIG_RCU_LAZY;
    - other.
- when started, invoking path is interrupted due to:
    - time limit;
    - need_resched();
    - if limit is reached.
- where in a nocb list it is located;
- how fast previous callbacks completed;

Example:

1. On our embedded devices i can easily trigger the scenario when
it is a last in the list out of ~3600 callbacks:

<snip>
  <...>-29      [001] d..1. 21950.145313: rcu_batch_start: rcu_preempt CBs=3613 bl=28
...
  <...>-29      [001] ..... 21950.152578: rcu_invoke_callback: rcu_preempt rhp=00000000b2d6dee8 func=__free_vm_area_struct.cfi_jt
  <...>-29      [001] ..... 21950.152579: rcu_invoke_callback: rcu_preempt rhp=00000000a446f607 func=__free_vm_area_struct.cfi_jt
  <...>-29      [001] ..... 21950.152580: rcu_invoke_callback: rcu_preempt rhp=00000000a5cab03b func=__free_vm_area_struct.cfi_jt
  <...>-29      [001] ..... 21950.152581: rcu_invoke_callback: rcu_preempt rhp=0000000013b7e5ee func=__free_vm_area_struct.cfi_jt
  <...>-29      [001] ..... 21950.152582: rcu_invoke_callback: rcu_preempt rhp=000000000a8ca6f9 func=__free_vm_area_struct.cfi_jt
  <...>-29      [001] ..... 21950.152583: rcu_invoke_callback: rcu_preempt rhp=000000008f162ca8 func=wakeme_after_rcu.cfi_jt
  <...>-29      [001] d..1. 21950.152625: rcu_batch_end: rcu_preempt CBs-invoked=3612 idle=....
<snip>

2. We use cpuset/cgroup to classify tasks and assign them into
different cgroups. For example "backgrond" group which binds tasks
only to little CPUs or "foreground" which makes use of all CPUs.
Tasks can be migrated between groups by a request if an acceleration
is needed.

See below an example how "surfaceflinger" task gets migrated.
Initially it is located in the "system-background" cgroup which
allows to run only on little cores. In order to speed it up it
can be temporary moved into "foreground" cgroup which allows
to use big/all CPUs:

cgroup_attach_task():
 -> cgroup_migrate_execute()
   -> cpuset_can_attach()
     -> percpu_down_write()
       -> rcu_sync_enter()
         -> synchronize_rcu()
   -> now move tasks to the new cgroup.
 -> cgroup_migrate_finish()

<snip>
         rcuop/1-29      [000] .....  7030.528570: rcu_invoke_callback: rcu_preempt rhp=00000000461605e0 func=wakeme_after_rcu.cfi_jt
    PERFD-SERVER-1855    [000] d..1.  7030.530293: cgroup_attach_task: dst_root=3 dst_id=22 dst_level=1 dst_path=/foreground pid=1900 comm=surfaceflinger
   TimerDispatch-2768    [002] d..5.  7030.537542: sched_migrate_task: comm=surfaceflinger pid=1900 prio=98 orig_cpu=0 dest_cpu=4
<snip>

"Boosting a task" depends on synchronize_rcu() latency:

- first trace shows a completion of synchronize_rcu();
- second shows attaching a task to a new group;
- last shows a final step when migration occurs.

3. To address this drawback, maintain a separate track that consists
of synchronize_rcu() callers only. After completion of a grace period
users are deferred to a dedicated worker to process requests.

4. This patch reduces the latency of synchronize_rcu() approximately
by ~30-40% on synthetic tests. The real test case, camera launch time,
shows(time is in milliseconds):

1-run 542 vs 489 improvement 9%
2-run 540 vs 466 improvement 13%
3-run 518 vs 468 improvement 9%
4-run 531 vs 457 improvement 13%
5-run 548 vs 475 improvement 13%
6-run 509 vs 484 improvement 4%

Synthetic test(no "noise" from other callbacks):
Hardware: x86_64 64 CPUs, 64GB of memory
Linux-6.6

- 10K tasks(simultaneous);
- each task does(1000 loops)
     synchronize_rcu();
     kfree(p);

default: CONFIG_RCU_NOCB_CPU: takes 54 seconds to complete all users;
patch: CONFIG_RCU_NOCB_CPU: takes 35 seconds to complete all users.

Running 60K gives approximately same results on my setup. Please note
it is without any interaction with another type of callbacks, otherwise
it will impact a lot a default case.

Signed-off-by: Uladzislau Rezki (Sony) <urezki@gmail.com>
---
 kernel/rcu/tree.c     | 135 +++++++++++++++++++++++++++++++++++++++++-
 kernel/rcu/tree_exp.h |   2 +-
 2 files changed, 135 insertions(+), 2 deletions(-)

Message ID	20231128080033.288050-2-urezki@gmail.com (mailing list archive)
State	New, archived
Headers	show Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Uz+j2EkX" Received: from mail-lf1-x135.google.com (mail-lf1-x135.google.com [IPv6:2a00:1450:4864:20::135]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 0CBD0CB; Tue, 28 Nov 2023 00:00:39 -0800 (PST) Received: by mail-lf1-x135.google.com with SMTP id 2adb3069b0e04-50bbb78efb5so372898e87.3; Tue, 28 Nov 2023 00:00:38 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1701158437; x=1701763237; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=6BsRDUs/wYt2ci1VcBpz4SJheODmui3/KELJxnudBXg=; b=Uz+j2EkXwZk+5VMMOPsa7J0g4mavGHnsCvbsedTmnUrH/vxUzWTMGWqANSJafcx3u9 bwvrcMbbb9ceauPpBrg+g8taVx9FnnbxPj2/ybM1g03ciE/TRAv/hQVNww29GcAo7Pmg dyqU+Wthq/ecL0EorDRTn2XNM+/5KrowHfb5ryGTR7//lHddGDQKF8v3bDIBiOeEPD/z ZxLB6kqz8hVFjkXUuNGHOXNB2BFH6LJ1AB6bI3uNQNdJhm+WWjnsIKgtU4lPW6FqqJyj xcz1q7pX3W2kqxRGRtxUOWCyRgXZN/pqwhM3CTjOPb3L4StIkbluExx4bsnnalwfmAqC V3+A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1701158437; x=1701763237; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=6BsRDUs/wYt2ci1VcBpz4SJheODmui3/KELJxnudBXg=; b=AlIl8NdiKm4zRKK+8ZaJP1i4Nh+k3xTR0nhBHDRI+KXvWa85EGWgefEbxgzS6XEM+3 3a75aASRg5NpHKE1mY0Wfm3t9rbN94V15ULKMsLLzD4RiJrL2bXZcGjFdlvnjFa1fy9R ju1ZtUQG6VZbRAnzCkqfCLnTx5OfC/aXw9kvtmxXyWOjIb2OtmkM4kC03OCvIYjULQK6 43QNorekJ7bGe+B76ldkIFEprufBPtzzlxpgT+EkMorV7WSvvd1nUWsMithi2Z/GsIQh TBDG0ycmXepL7GqMEecPTzxwqyEUKvO1vsA6EvPPuLdzYEmCYP/bi/rel7XnzE9VR+Ie GREA== X-Gm-Message-State: AOJu0YzM/S5AB42V6F/g22+1UG3qlMWpR7DFMOr1ShUcC4y2o+/TyQHd Yzpl5isrwzVqf2eo+4XkzSs= X-Google-Smtp-Source: AGHT+IH9AKjFYiSSwd4Gl/X8xCmeiuE6ev0b1jilePlkMzKYVASLtbE6S3FbybnDTVZn4LqRpYskDw== X-Received: by 2002:a05:6512:38cf:b0:503:3781:ac32 with SMTP id p15-20020a05651238cf00b005033781ac32mr8391110lft.41.1701158436996; Tue, 28 Nov 2023 00:00:36 -0800 (PST) Received: from pc638.lan ([155.137.26.201]) by smtp.gmail.com with ESMTPSA id o16-20020ac24bd0000000b004fe202a5c7csm1765501lfq.135.2023.11.28.00.00.36 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 28 Nov 2023 00:00:36 -0800 (PST) From: "Uladzislau Rezki (Sony)" <urezki@gmail.com> To: "Paul E . McKenney" <paulmck@kernel.org> Cc: RCU <rcu@vger.kernel.org>, Neeraj upadhyay <Neeraj.Upadhyay@amd.com>, Boqun Feng <boqun.feng@gmail.com>, Hillf Danton <hdanton@sina.com>, Joel Fernandes <joel@joelfernandes.org>, LKML <linux-kernel@vger.kernel.org>, Uladzislau Rezki <urezki@gmail.com>, Oleksiy Avramchenko <oleksiy.avramchenko@sony.com>, Frederic Weisbecker <frederic@kernel.org> Subject: [PATCH v3 1/7] rcu: Reduce synchronize_rcu() latency Date: Tue, 28 Nov 2023 09:00:27 +0100 Message-Id: <20231128080033.288050-2-urezki@gmail.com> X-Mailer: git-send-email 2.39.2 In-Reply-To: <20231128080033.288050-1-urezki@gmail.com> References: <20231128080033.288050-1-urezki@gmail.com> Precedence: bulk X-Mailing-List: rcu@vger.kernel.org List-Id: <rcu.vger.kernel.org> List-Subscribe: <mailto:rcu+subscribe@vger.kernel.org> List-Unsubscribe: <mailto:rcu+unsubscribe@vger.kernel.org> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit
Series	Reduce synchronize_rcu() latency(V3) \| expand [v3,0/7] Reduce synchronize_rcu() latency(V3) [v3,1/7] rcu: Reduce synchronize_rcu() latency [v3,2/7] rcu: Add a trace event for synchronize_rcu_normal() [v3,3/7] doc: Add rcutree.rcu_normal_wake_from_gp to kernel-parameters.txt [v3,4/7] rcu: Improve handling of synchronize_rcu() users [v3,5/7] rcu: Support direct wake-up of synchronize_rcu() users [v3,6/7] rcu: Move sync related data to rcu_state structure [v3,7/7] rcu: Add CONFIG_RCU_SR_NORMAL_DEBUG_GP

[v3,1/7] rcu: Reduce synchronize_rcu() latency

Commit Message

Patch