From patchwork Tue Nov 21 17:16:34 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Suren Baghdasaryan X-Patchwork-Id: 13463393 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 3375AC61D92 for ; Tue, 21 Nov 2023 17:16:53 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 952986B04C3; Tue, 21 Nov 2023 12:16:52 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id 901606B04C4; Tue, 21 Nov 2023 12:16:52 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 7328D6B04C5; Tue, 21 Nov 2023 12:16:52 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0014.hostedemail.com [216.40.44.14]) by kanga.kvack.org (Postfix) with ESMTP id 5F0516B04C3 for ; Tue, 21 Nov 2023 12:16:52 -0500 (EST) Received: from smtpin19.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay04.hostedemail.com (Postfix) with ESMTP id 2DE0F1A051D for ; Tue, 21 Nov 2023 17:16:52 +0000 (UTC) X-FDA: 81482616264.19.033D17D Received: from mail-yw1-f201.google.com (mail-yw1-f201.google.com [209.85.128.201]) by imf15.hostedemail.com (Postfix) with ESMTP id 4D5DAA0023 for ; Tue, 21 Nov 2023 17:16:50 +0000 (UTC) Authentication-Results: imf15.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=ZFIBBeWo; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf15.hostedemail.com: domain of 3AeZcZQYKCHIikhUdRWeeWbU.SecbYdkn-ccalQSa.ehW@flex--surenb.bounces.google.com designates 209.85.128.201 as permitted sender) smtp.mailfrom=3AeZcZQYKCHIikhUdRWeeWbU.SecbYdkn-ccalQSa.ehW@flex--surenb.bounces.google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1700587010; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=He74YIcSIwOjGCEm7bMpO/TfFvRAcUHKo6jbQtwLQrs=; b=tJFBhmNNhugquOxgvkzVpgtlyRfCnW2MZigEKys0xHbF/1HpjsIYPtg40J2GnfBqTxYq5D Eyg5GjcIDm85Z9xMTOOIrjVaw/W8cv9p5W+4TvZiTgzIMPkrkr3EYELaN6ZZhRXB9iCdns 998Gw0hTYVbCvrSa30tOGBF1xIwcqig= ARC-Authentication-Results: i=1; imf15.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=ZFIBBeWo; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf15.hostedemail.com: domain of 3AeZcZQYKCHIikhUdRWeeWbU.SecbYdkn-ccalQSa.ehW@flex--surenb.bounces.google.com designates 209.85.128.201 as permitted sender) smtp.mailfrom=3AeZcZQYKCHIikhUdRWeeWbU.SecbYdkn-ccalQSa.ehW@flex--surenb.bounces.google.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700587010; a=rsa-sha256; cv=none; b=q9GCR6kJm5voSMRpEyWptgISJttGHViwZxhdjc3z/tq+jm2v0U2p23qPFSvge6S42SOBpK 4jB0zKomvSR5VyNsCnFvsRXtY7P2t+aciQSyV/soeG47Vun8e8rsKCjKH2hTf0v0OGE0mg WDxYAatDJszNpkQRKIyxZHL8ZoAC2RY= Received: by mail-yw1-f201.google.com with SMTP id 00721157ae682-5c87663a873so62597987b3.2 for ; Tue, 21 Nov 2023 09:16:50 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1700587009; x=1701191809; darn=kvack.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=He74YIcSIwOjGCEm7bMpO/TfFvRAcUHKo6jbQtwLQrs=; b=ZFIBBeWoIu/4KUmyf3dXBwDV1j/AwzL7Mysee1o3Ha8Zhliv4/vLqNnLuX4PaxJraI svPDEfiDrXes8BZNU+tjP9buvPnZbZ+6Tiz6+F8Sy0PQojyX0RU1aS7nZ5bUX88+E5B7 HgvoKOLZVLTwDplsxxb5hJnfXhjgp1UYnhRhtXbIporNzISMhLkp9ookzF1PUkm/yXAe ODyawm0e3ADD/I6CdzXvbD7EFqE9gJpK0RM0HZDJrggCZgcMXimcERRcB/3pUhbaIdYB i9sf+rMZjLY5CKqgXw7rZc0/VZQhl+D/26QiV7FV2O88GjGip/ytJ+kowCDmMXMQOXve b5nQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1700587009; x=1701191809; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=He74YIcSIwOjGCEm7bMpO/TfFvRAcUHKo6jbQtwLQrs=; b=Yk57jaDVsEunYTn+tY+hPfBnjxTcaKLnqH99TSBLkvBx5xAwDM+wTxfcGueQD/Ld1r uAJJRks3Hfp5mmNeAJaHGNkmvnuuiuBp6LkAJGATH5txe+oSZp5kW2sXIBRTB2OLcWKw +9pccugCGqNJ1ZFJnXV3sK5IWTDFr2RAFXZQ7sdnEBPFdqpSrdrAAqPGVY/Rap2uTjBI gfqX38NyOh19+SfrDz7mzq135spwBfJz1yzc+ITj4jRS4CUBaO6u10EsGIBfTfWuSrDT zi39A4DcJXXwgz87ROUSPkGMq0fjGquOvkC4CPR2NPkvLUX/pjPvAP3auXlzpbvKKx8o g8qA== X-Gm-Message-State: AOJu0Yy58RQl1YL5wdp9kKPfgWbj8BMzhS2ZMenntEiynS7KzvesXtZ0 uq2jQVXC7tHxAvglN/WABMDEWe7ALEQ= X-Google-Smtp-Source: AGHT+IGyUrMT5Mj/g7lBAQcYKLbTTgQRTtXQcXK7X2hl+2vA5HYqDMy63skEG09MWnl+1v4a8CZ5k0TbKFY= X-Received: from surenb-desktop.mtv.corp.google.com ([2620:15c:211:201:2045:f6d2:f01d:3fff]) (user=surenb job=sendgmr) by 2002:a0d:fb03:0:b0:5c8:b756:f3af with SMTP id l3-20020a0dfb03000000b005c8b756f3afmr330699ywf.4.1700587009501; Tue, 21 Nov 2023 09:16:49 -0800 (PST) Date: Tue, 21 Nov 2023 09:16:34 -0800 In-Reply-To: <20231121171643.3719880-1-surenb@google.com> Mime-Version: 1.0 References: <20231121171643.3719880-1-surenb@google.com> X-Mailer: git-send-email 2.43.0.rc1.413.gea7ed67945-goog Message-ID: <20231121171643.3719880-2-surenb@google.com> Subject: [PATCH v5 1/5] mm/rmap: support move to different root anon_vma in folio_move_anon_rmap() From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: viro@zeniv.linux.org.uk, brauner@kernel.org, shuah@kernel.org, aarcange@redhat.com, lokeshgidra@google.com, peterx@redhat.com, david@redhat.com, hughd@google.com, mhocko@suse.com, axelrasmussen@google.com, rppt@kernel.org, willy@infradead.org, Liam.Howlett@oracle.com, jannh@google.com, zhangpeng362@huawei.com, bgeffon@google.com, kaleshsingh@google.com, ngeoffray@google.com, jdduke@google.com, surenb@google.com, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, kernel-team@android.com X-Rspamd-Queue-Id: 4D5DAA0023 X-Rspam-User: X-Rspamd-Server: rspam04 X-Stat-Signature: df53x61yre399wawahjzwd1h4xabhy3b X-HE-Tag: 1700587010-151162 X-HE-Meta: U2FsdGVkX1+RLbU+GM1z9WpqjrWyFg7HHA7k6yVMKa5kEBGAXzF1Woj83/LL8uSf7ZD1I8jnk1MCs1qbhYpkdPkudbRU09vWQml4PhPLixm2XdQuQDcPh1ILJFNTqCk1czNKWLGKloOTum6DbPLGBdSpCjDLy7pfXpa9OeMHuXzTRpgTCJGG0+0Xipt4sltVHazRayOiLV8ZED7YhHOb5eVA6bu/WPPFw5WdVahEHA9zOp88A5XxBsdD4hm188HlMCN0N35fO2TcUlNBc3NzJxDPqsZQjX52o8c5D96gnUmV3n2az5iveI+g057s0OQCbI4QOijb5e4Ed4iGt4LSbFx4xQp/68gQBYu3lQLKA9tYI4ISXPZsu/STQgqhn5lEfPuNAGQfTatZr3TCurdnrAAWf/jx4h/V8LItmqikMihhIu+Hq4+sSNjA7qeF4VUQ1gl8LmoxJf5EipMtrEV18R9iDQah2Se9eaPjirPHoUONl+FOSxQ9FPpYpY5nZqb6EVbUTaftNwm1m+fLcG8onXAGx0JrNtSPFNdeG65aCpVVgF6kGvsSR2cxMOPBiyOrvL2CJCmejLZPuKau0K0nA95imsqsGl8NLwCCwKvDpsTlzNt5JLz+muEXtSXcvv4ddOgHTyAGNZ0YCKvOeL43D/gRuc6zonKBXJ2SbcSppaTquyvRLxu+99kSpBz0LaLTgTYZ8pVxxn0SvFQNaMgLF5r+8p6RNUWpEfG/9wOeX0EiNDCJntMJl2U7W8T23JaRFp7pGxYaVqYj1Y6fAXO6KHaO/51DBW2rVziglCcmTCdaBE29SBxVc5aiFe6Fy/dWuLaPjs3UCdYzG9esUe4ELY6qEsMh0MBpqGK7FhEPFL+iPKIJLzmIKkbyO0Wb3NbtSQr1MkYJVnVy/bSil6AhgjM6m7gkjlzfzUMHABUymLU7bO/6SCmtEHftyI+bvMkd/OCK8/0Nwy0rZNqc2Sn VKkNXA3q C9h1igehND4j7++ag6h8g2kgCGaFg6XaBJOwbT1ATQ/Gjxcw3nvapzTNVFsmjfi1e8t3thylPdXWL952pjuYiZL84m1YM+I75feWlI0kWZhLJoX3KPRNQLhk7rgkHzwdGz72UBUf670nKQu6k57Ipx+LSW19u9Vk1zRTdsDVdvWUlLBpxWVq5Q2d0FNFp4zYNKI4bmx1C4yCmSjUEGUgl0ToiSQicfn3IaBaRbmUdQ1waK1urgGjJFPILp56wk90tbijT6Z4KQl5Esn//q803iHdJpqc9JYCZ20s6YK5FVddiN9/3GAPOHaZeyfpXtrLANHX2QpdPCjYY9ZxUW9iM+6tHhAXI1F+D8SVSdSQpUBh21V6a3MePTGtW8APnWBwMJq37RyWtVQRzo5YbLLawq/7hPRdEyq5AydvSdbJmACyjue9tGYZIAm7GX5Zzf92VhMenSCRpZvncc/JzYVIwPJeLyNyOQ1PTWDYBlLLN8N9l4hWf2woA+Zu9FgGd+xHmWkCqUqFcgVNtIepNQR2/3VkCRfPNK6xXI07O6u+6NiD6+hhB3Uuatyv+3s1yxR98MWGjTIXMLy8PV2EZqnKiXVKcNJZD/19SXQCgpsV4JS+XpN+aUAh/dLptP3n8ZZPfQuSgGOaL92pap4aWn46rYlaHQFKIPOiDKHLBjIFzTfEXDTlaFjDV/XYvSg== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: From: Andrea Arcangeli For now, folio_move_anon_rmap() was only used to move a folio to a different anon_vma after fork(), whereby the root anon_vma stayed unchanged. For that, it was sufficient to hold the folio lock when calling folio_move_anon_rmap(). However, we want to make use of folio_move_anon_rmap() to move folios between VMAs that have a different root anon_vma. As folio_referenced() performs an RMAP walk without holding the folio lock but only holding the anon_vma in read mode, holding the folio lock is insufficient. When moving to an anon_vma with a different root anon_vma, we'll have to hold both, the folio lock and the anon_vma lock in write mode. Consequently, whenever we succeeded in folio_lock_anon_vma_read() to read-lock the anon_vma, we have to re-check if the mapping was changed in the meantime. If that was the case, we have to retry. Note that folio_move_anon_rmap() must only be called if the anon page is exclusive to a process, and must not be called on KSM folios. This is a preparation for UFFDIO_MOVE, which will hold the folio lock, the anon_vma lock in write mode, and the mmap_lock in read mode. Signed-off-by: Andrea Arcangeli Signed-off-by: Suren Baghdasaryan Acked-by: Peter Xu --- mm/rmap.c | 24 ++++++++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/mm/rmap.c b/mm/rmap.c index 7a27a2b41802..525c5bc0b0b3 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -542,6 +542,7 @@ struct anon_vma *folio_lock_anon_vma_read(struct folio *folio, struct anon_vma *root_anon_vma; unsigned long anon_mapping; +retry: rcu_read_lock(); anon_mapping = (unsigned long)READ_ONCE(folio->mapping); if ((anon_mapping & PAGE_MAPPING_FLAGS) != PAGE_MAPPING_ANON) @@ -552,6 +553,17 @@ struct anon_vma *folio_lock_anon_vma_read(struct folio *folio, anon_vma = (struct anon_vma *) (anon_mapping - PAGE_MAPPING_ANON); root_anon_vma = READ_ONCE(anon_vma->root); if (down_read_trylock(&root_anon_vma->rwsem)) { + /* + * folio_move_anon_rmap() might have changed the anon_vma as we + * might not hold the folio lock here. + */ + if (unlikely((unsigned long)READ_ONCE(folio->mapping) != + anon_mapping)) { + up_read(&root_anon_vma->rwsem); + rcu_read_unlock(); + goto retry; + } + /* * If the folio is still mapped, then this anon_vma is still * its anon_vma, and holding the mutex ensures that it will @@ -586,6 +598,18 @@ struct anon_vma *folio_lock_anon_vma_read(struct folio *folio, rcu_read_unlock(); anon_vma_lock_read(anon_vma); + /* + * folio_move_anon_rmap() might have changed the anon_vma as we might + * not hold the folio lock here. + */ + if (unlikely((unsigned long)READ_ONCE(folio->mapping) != + anon_mapping)) { + anon_vma_unlock_read(anon_vma); + put_anon_vma(anon_vma); + anon_vma = NULL; + goto retry; + } + if (atomic_dec_and_test(&anon_vma->refcount)) { /* * Oops, we held the last refcount, release the lock From patchwork Tue Nov 21 17:16:35 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Patchwork-Submitter: Suren Baghdasaryan X-Patchwork-Id: 13463394 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id CD8CFC61D96 for ; Tue, 21 Nov 2023 17:16:55 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 50CBB6B04C5; Tue, 21 Nov 2023 12:16:55 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id 4971F6B04C6; Tue, 21 Nov 2023 12:16:55 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 27E1C6B04C9; Tue, 21 Nov 2023 12:16:55 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0014.hostedemail.com [216.40.44.14]) by kanga.kvack.org (Postfix) with ESMTP id 0F34E6B04C5 for ; Tue, 21 Nov 2023 12:16:55 -0500 (EST) Received: from smtpin08.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay06.hostedemail.com (Postfix) with ESMTP id D0A62B5CC2 for ; Tue, 21 Nov 2023 17:16:54 +0000 (UTC) X-FDA: 81482616348.08.87FD3E7 Received: from mail-yb1-f202.google.com (mail-yb1-f202.google.com [209.85.219.202]) by imf21.hostedemail.com (Postfix) with ESMTP id B45801C0025 for ; Tue, 21 Nov 2023 17:16:52 +0000 (UTC) Authentication-Results: imf21.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=hyQvxHv8; spf=pass (imf21.hostedemail.com: domain of 3A-ZcZQYKCHQkmjWfTYggYdW.Ugedafmp-eecnSUc.gjY@flex--surenb.bounces.google.com designates 209.85.219.202 as permitted sender) smtp.mailfrom=3A-ZcZQYKCHQkmjWfTYggYdW.Ugedafmp-eecnSUc.gjY@flex--surenb.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1700587012; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=31fA93BiuMj4g20IEIp4qZ8JUc+CMYYBAAS6MShowx0=; b=Fjlp8SE58eLW8PDT4dhKqosGBjSG+SyMrRcQ2uw6F6lPxzuSEfmGR7aCOFVbEskxBw1+aK kYdTLrETB0/Kha5S1t8K2ByviOhP3rfxYEtxQOYMNhS6cJFTUg/Ww9BCEg4a72HkQz44y+ 5ZPOP+xg+cdHs8RHMm5exktzddP8yOg= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700587012; a=rsa-sha256; cv=none; b=BYIs2hwe6ikxe9TkHe+j+Ov8JaDij32t7oGxqhfMCRha8FTboORpgyVNcAUaycyRFkbgrg ScbHHjjbn5ZZR9161jFW1qV8pqJ9oE/1E4x/feGerC286dM7boqmq3l62QuqmvPXjY2m7v sk2S1uYWTQAgxGkVTZ3DHduR29vT414= ARC-Authentication-Results: i=1; imf21.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=hyQvxHv8; spf=pass (imf21.hostedemail.com: domain of 3A-ZcZQYKCHQkmjWfTYggYdW.Ugedafmp-eecnSUc.gjY@flex--surenb.bounces.google.com designates 209.85.219.202 as permitted sender) smtp.mailfrom=3A-ZcZQYKCHQkmjWfTYggYdW.Ugedafmp-eecnSUc.gjY@flex--surenb.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com Received: by mail-yb1-f202.google.com with SMTP id 3f1490d57ef6-daee86e2d70so7076240276.0 for ; Tue, 21 Nov 2023 09:16:52 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1700587012; x=1701191812; darn=kvack.org; h=content-transfer-encoding:cc:to:from:subject:message-id:references :mime-version:in-reply-to:date:from:to:cc:subject:date:message-id :reply-to; bh=31fA93BiuMj4g20IEIp4qZ8JUc+CMYYBAAS6MShowx0=; b=hyQvxHv8f9Xe88cFJGqarLrLfORq03SkJkDdH0gI+28jFxUquj41KXfWyGZGBhG1q6 fl6rk/EAT1MLFEmv14CRJiz+z5/euoIOKHFNJFxXR2eQpRhklVOEukTnh1cb/2CLBWNg hgX8nqGan5pkRCVnixw3yPVlZRZ0BpuTC+4kf7+qy3VSnUAt4f1JB62Oou9DPB38168O DjfqiCLZaU2Ay3bDH2if610fw+EX0sFNPUS6kKmBXvgpOm50+u5rsyee9pNToETjopZO 6MiNP5wGTMhRs7/kUpCg8bG+EOAro3UUDn/GfgY01kvaNz/OVKsXud5Ca5K6sFBrPyia 8JIA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1700587012; x=1701191812; h=content-transfer-encoding:cc:to:from:subject:message-id:references :mime-version:in-reply-to:date:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to; bh=31fA93BiuMj4g20IEIp4qZ8JUc+CMYYBAAS6MShowx0=; b=wPpSPLbL1+hXGVKEvUrYmc+Uuc5lS7VkAw0I3ZRaBBO3LpWok9i5jaSp0oMsE8RF1Q KefXkJ6MsLCYcs418paXXVvDqmfJUwDfQJ5Ky5LNvA7iCcObEtTW3R5qplC6bQfn2Elo 7lIIJSg6gxlDkXIo7Pm6mjU4iHvg9NNlb+l/tCRk43pNAa0660mqE0TrU/hafJP0Vu1t AewKhyouMJpSV+7FSk4L2ThwWZKlngOcjdfHOyJGmV/E7KcyKQm7jcVKJLam4GtmBAyy VLbOFxHf+G4VhPwYySf4GWhAN7TK6M9xvuul3vXQgWqgfR75F0V8G/xDIOJQElqvrJZ4 dBaw== X-Gm-Message-State: AOJu0Yx4n3KI18JmFmVcAxblCKrEPkulB1mwaVVvvU7HUqnTikQcghih 3KHYdT75W7gvV1TGywfTEy6H+CS2tg8= X-Google-Smtp-Source: AGHT+IE08o6eao7b7LQK+YNaqVMpIRR7JC4+TfHhBYaKmqDVmkGyMGNB75c4qvQf06tGjwlxnWEa6gYeHJI= X-Received: from surenb-desktop.mtv.corp.google.com ([2620:15c:211:201:2045:f6d2:f01d:3fff]) (user=surenb job=sendgmr) by 2002:a25:25d6:0:b0:daf:7949:52ed with SMTP id l205-20020a2525d6000000b00daf794952edmr235129ybl.8.1700587011814; Tue, 21 Nov 2023 09:16:51 -0800 (PST) Date: Tue, 21 Nov 2023 09:16:35 -0800 In-Reply-To: <20231121171643.3719880-1-surenb@google.com> Mime-Version: 1.0 References: <20231121171643.3719880-1-surenb@google.com> X-Mailer: git-send-email 2.43.0.rc1.413.gea7ed67945-goog Message-ID: <20231121171643.3719880-3-surenb@google.com> Subject: [PATCH v5 2/5] userfaultfd: UFFDIO_MOVE uABI From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: viro@zeniv.linux.org.uk, brauner@kernel.org, shuah@kernel.org, aarcange@redhat.com, lokeshgidra@google.com, peterx@redhat.com, david@redhat.com, hughd@google.com, mhocko@suse.com, axelrasmussen@google.com, rppt@kernel.org, willy@infradead.org, Liam.Howlett@oracle.com, jannh@google.com, zhangpeng362@huawei.com, bgeffon@google.com, kaleshsingh@google.com, ngeoffray@google.com, jdduke@google.com, surenb@google.com, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, kernel-team@android.com X-Rspamd-Queue-Id: B45801C0025 X-Rspam-User: X-Rspamd-Server: rspam11 X-Stat-Signature: ee81qtoyezbwzoxenzsuh4sumw5t7z7r X-HE-Tag: 1700587012-682483 X-HE-Meta: U2FsdGVkX1/BXy4UeA+0PjJ6H6L+1Bh6qbAmf79uzs/i97AtAMHHD1URuca2qjTpi9kKa+/PAauEFZfEsH7n8Yxiwm/+/0G0m+O/WOncPBOgXVWkehiJsOuLQpSfJwGkzh5n9PSto8Ifc6isoPwyVen250E9B0dfaGh35I4MK4JwkVdldmZWGdTAqqB4c0L/z0FX3ABLivi8O6px6hSGa+l8xXga/BPrP27t1OwMlnniClZEaLVxPwLWtZI7myq9/r89QV27LfvmRwnirnIMYLMOuoRqnzmEULyG2LR9Jil8NM1zXkYZ8yKoW+enQHqToWtkdTQvn++TBaTqV/LhDN3rXIveEOZTkH4LFCR9CwEEMN+wZsoiMBkHMlZH476RWVjjpOMXMLAsxRbr0fRk+zmQ28Jw8ifRsqw3zf1fTzuj4NWi1zn32rhJJWJlGi759wxOgCFpqC/WdvpZLEL14s4ahBnoFNqvSAaEPIEDrRo4JGs8h7Bqn9P7Ck5nHF7ofmGhyazXqOhpnurpfShHc4CWxKUxEB1epMRDJsmYsF3SWJceDTgHyseYmSsGDwFmExXOVu+Wud5DkXVUlUtEOJdNDTniQ1GdZm0WZFB4Ewr87a1ptloK79oJIyXfRPKDlE16cZyenLsvnRG2/q6yItm+FQj/3pMLxr81DFipHWtKTViT0+If492iD3sVg2FQhKiOcXPJuQ+/7tU+r9CTuD8HKdK88ivf7b0ujFjs9wnsDNc4knfovPprLWTZYaiRvYpM8CohtaUSycxU++FzhDQupaq1b+T7vDr6gEXzXg1NoU1v5FHLC42P9czItY398lTDCZH1tRBfkaMTSX8O6NcmMCxhiDwpCaBzWdHrhxKtnsck/OWqS9rZ51XIR5HhAAe4CI0SJSqsEzfGP6HRWKZP6jpSnUC8hIHMKGHHkjibArMHMboRcPM6bjJlT92E+ahSqlA+9NoR/nv+Kjn wRCkMP0Y LeSIJhCq/jlZ+f9e0FW5S53g0uP7XE74rOkVUUNciFEeBK8Cwpj/o8BZGIjVRJlu+GaC8hA9aIsrzeu4LhO2mZxt3fGXaimBFWkfU9bzdduyeKe8B71UCXjFwML45yoQOEJuGmC3aDUmSHS0Blt14SQfIuZ74ykIIkGy7xcwOjQw5uL7IHT34Yu6Kk/0sGEFyZdORVyDat4yskJiG+2Qx5cu8EuqbrXARufTpLoBXwgyIxcH+hx3ODuvKS33plgSgXw0pwaJE6pWcbR8hZFORz0zyaVlkprxdcYenxXwL3MNaWnfe9cyY+7fo04Uo0QjYvbzdaVb8QPvtXX5FZ0yBwk2bM2Bm94IdUGsC9vUemi7gYGVgcFtsMvXHLbt1hMPdBi1AOP04E5mDUkEpirFd4ypPF8HebS678rU/YoQDvwJCfj39PjUtPS0lI0lHaopWBJOj0rf/BAcq+ayXxwSrbFgoCQl6W6HpxwjcyVwtVyVwJOm9+QxHjBAdD5QXXu4+WsIt8GWH4EF955IDb2AA1rNaq4GBSX88VB/HCRRHACIzTeW0V9RrC0DmBTQ289Z4CqNZjBrlyxGMbniLFoLTPitXED7Q3OzarZz3Vz5p+HUpQzPy9CrY1ldksYSHxTfSXZqDkvI1tLZ1VTyURYkD82Q1gqlmHTGjIN6mJl+jmBMy+rBO/WmquRH6VW5y7/AdBhILst3yq+q231wCTmnTVBbT7Y7Wv2cRpr0xu0y9AUYc7Vr/5MOaQO0bKaAe8b5ZqiH3 X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: From: Andrea Arcangeli Implement the uABI of UFFDIO_MOVE ioctl. UFFDIO_COPY performs ~20% better than UFFDIO_MOVE when the application needs pages to be allocated [1]. However, with UFFDIO_MOVE, if pages are available (in userspace) for recycling, as is usually the case in heap compaction algorithms, then we can avoid the page allocation and memcpy (done by UFFDIO_COPY). Also, since the pages are recycled in the userspace, we avoid the need to release (via madvise) the pages back to the kernel [2]. We see over 40% reduction (on a Google pixel 6 device) in the compacting thread’s completion time by using UFFDIO_MOVE vs. UFFDIO_COPY. This was measured using a benchmark that emulates a heap compaction implementation using userfaultfd (to allow concurrent accesses by application threads). More details of the usecase are explained in [2]. Furthermore, UFFDIO_MOVE enables moving swapped-out pages without touching them within the same vma. Today, it can only be done by mremap, however it forces splitting the vma. [1] https://lore.kernel.org/all/1425575884-2574-1-git-send-email-aarcange@redhat.com/ [2] https://lore.kernel.org/linux-mm/CA+EESO4uO84SSnBhArH4HvLNhaUQ5nZKNKXqxRCyjniNVjp0Aw@mail.gmail.com/ Update for the ioctl_userfaultfd(2) manpage: UFFDIO_MOVE (Since Linux xxx) Move a continuous memory chunk into the userfault registered range and optionally wake up the blocked thread. The source and destination addresses and the number of bytes to move are specified by the src, dst, and len fields of the uffdio_move structure pointed to by argp: struct uffdio_move { __u64 dst; /* Destination of move */ __u64 src; /* Source of move */ __u64 len; /* Number of bytes to move */ __u64 mode; /* Flags controlling behavior of move */ __s64 move; /* Number of bytes moved, or negated error */ }; The following value may be bitwise ORed in mode to change the behavior of the UFFDIO_MOVE operation: UFFDIO_MOVE_MODE_DONTWAKE Do not wake up the thread that waits for page-fault resolution UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES Allow holes in the source virtual range that is being moved. When not specified, the holes will result in ENOENT error. When specified, the holes will be accounted as successfully moved memory. This is mostly useful to move hugepage aligned virtual regions without knowing if there are transparent hugepages in the regions or not, but preventing the risk of having to split the hugepage during the operation. The move field is used by the kernel to return the number of bytes that was actually moved, or an error (a negated errno- style value). If the value returned in move doesn't match the value that was specified in len, the operation fails with the error EAGAIN. The move field is output-only; it is not read by the UFFDIO_MOVE operation. The operation may fail for various reasons. Usually, remapping of pages that are not exclusive to the given process fail; once KSM might deduplicate pages or fork() COW-shares pages during fork() with child processes, they are no longer exclusive. Further, the kernel might only perform lightweight checks for detecting whether the pages are exclusive, and return -EBUSY in case that check fails. To make the operation more likely to succeed, KSM should be disabled, fork() should be avoided or MADV_DONTFORK should be configured for the source VMA before fork(). This ioctl(2) operation returns 0 on success. In this case, the entire area was moved. On error, -1 is returned and errno is set to indicate the error. Possible errors include: EAGAIN The number of bytes moved (i.e., the value returned in the move field) does not equal the value that was specified in the len field. EINVAL Either dst or len was not a multiple of the system page size, or the range specified by src and len or dst and len was invalid. EINVAL An invalid bit was specified in the mode field. ENOENT The source virtual memory range has unmapped holes and UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES is not set. EEXIST The destination virtual memory range is fully or partially mapped. EBUSY The pages in the source virtual memory range are either pinned or not exclusive to the process. The kernel might only perform lightweight checks for detecting whether the pages are exclusive. To make the operation more likely to succeed, KSM should be disabled, fork() should be avoided or MADV_DONTFORK should be configured for the source virtual memory area before fork(). ENOMEM Allocating memory needed for the operation failed. ESRCH The target process has exited at the time of a UFFDIO_MOVE operation. Signed-off-by: Andrea Arcangeli Signed-off-by: Suren Baghdasaryan --- Documentation/admin-guide/mm/userfaultfd.rst | 3 + fs/userfaultfd.c | 72 +++ include/linux/rmap.h | 5 + include/linux/userfaultfd_k.h | 11 + include/uapi/linux/userfaultfd.h | 29 +- mm/huge_memory.c | 122 ++++ mm/khugepaged.c | 3 + mm/rmap.c | 6 + mm/userfaultfd.c | 599 +++++++++++++++++++ 9 files changed, 849 insertions(+), 1 deletion(-) diff --git a/Documentation/admin-guide/mm/userfaultfd.rst b/Documentation/admin-guide/mm/userfaultfd.rst index 203e26da5f92..e5cc8848dcb3 100644 --- a/Documentation/admin-guide/mm/userfaultfd.rst +++ b/Documentation/admin-guide/mm/userfaultfd.rst @@ -113,6 +113,9 @@ events, except page fault notifications, may be generated: areas. ``UFFD_FEATURE_MINOR_SHMEM`` is the analogous feature indicating support for shmem virtual memory areas. +- ``UFFD_FEATURE_MOVE`` indicates that the kernel supports moving an + existing page contents from userspace. + The userland application should set the feature flags it intends to use when invoking the ``UFFDIO_API`` ioctl, to request that those features be enabled if supported. diff --git a/fs/userfaultfd.c b/fs/userfaultfd.c index e8af40b05549..6e2a4d6a0d8f 100644 --- a/fs/userfaultfd.c +++ b/fs/userfaultfd.c @@ -2005,6 +2005,75 @@ static inline unsigned int uffd_ctx_features(__u64 user_features) return (unsigned int)user_features | UFFD_FEATURE_INITIALIZED; } +static int userfaultfd_move(struct userfaultfd_ctx *ctx, + unsigned long arg) +{ + __s64 ret; + struct uffdio_move uffdio_move; + struct uffdio_move __user *user_uffdio_move; + struct userfaultfd_wake_range range; + struct mm_struct *mm = ctx->mm; + + user_uffdio_move = (struct uffdio_move __user *) arg; + + if (atomic_read(&ctx->mmap_changing)) + return -EAGAIN; + + if (copy_from_user(&uffdio_move, user_uffdio_move, + /* don't copy "move" last field */ + sizeof(uffdio_move)-sizeof(__s64))) + return -EFAULT; + + /* Do not allow cross-mm moves. */ + if (mm != current->mm) + return -EINVAL; + + ret = validate_range(mm, uffdio_move.dst, uffdio_move.len); + if (ret) + return ret; + + ret = validate_range(mm, uffdio_move.src, uffdio_move.len); + if (ret) + return ret; + + if (uffdio_move.mode & ~(UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES| + UFFDIO_MOVE_MODE_DONTWAKE)) + return -EINVAL; + + if (mmget_not_zero(mm)) { + mmap_read_lock(mm); + + /* Re-check after taking mmap_lock */ + if (likely(!atomic_read(&ctx->mmap_changing))) + ret = move_pages(ctx, mm, uffdio_move.dst, uffdio_move.src, + uffdio_move.len, uffdio_move.mode); + else + ret = -EINVAL; + + mmap_read_unlock(mm); + mmput(mm); + } else { + return -ESRCH; + } + + if (unlikely(put_user(ret, &user_uffdio_move->move))) + return -EFAULT; + if (ret < 0) + goto out; + + /* len == 0 would wake all */ + VM_WARN_ON(!ret); + range.len = ret; + if (!(uffdio_move.mode & UFFDIO_MOVE_MODE_DONTWAKE)) { + range.start = uffdio_move.dst; + wake_userfault(ctx, &range); + } + ret = range.len == uffdio_move.len ? 0 : -EAGAIN; + +out: + return ret; +} + /* * userland asks for a certain API version and we return which bits * and ioctl commands are implemented in this kernel for such API @@ -2097,6 +2166,9 @@ static long userfaultfd_ioctl(struct file *file, unsigned cmd, case UFFDIO_ZEROPAGE: ret = userfaultfd_zeropage(ctx, arg); break; + case UFFDIO_MOVE: + ret = userfaultfd_move(ctx, arg); + break; case UFFDIO_WRITEPROTECT: ret = userfaultfd_writeprotect(ctx, arg); break; diff --git a/include/linux/rmap.h b/include/linux/rmap.h index b26fe858fd44..8034eda972e5 100644 --- a/include/linux/rmap.h +++ b/include/linux/rmap.h @@ -121,6 +121,11 @@ static inline void anon_vma_lock_write(struct anon_vma *anon_vma) down_write(&anon_vma->root->rwsem); } +static inline int anon_vma_trylock_write(struct anon_vma *anon_vma) +{ + return down_write_trylock(&anon_vma->root->rwsem); +} + static inline void anon_vma_unlock_write(struct anon_vma *anon_vma) { up_write(&anon_vma->root->rwsem); diff --git a/include/linux/userfaultfd_k.h b/include/linux/userfaultfd_k.h index f2dc19f40d05..e4056547fbe6 100644 --- a/include/linux/userfaultfd_k.h +++ b/include/linux/userfaultfd_k.h @@ -93,6 +93,17 @@ extern int mwriteprotect_range(struct mm_struct *dst_mm, extern long uffd_wp_range(struct vm_area_struct *vma, unsigned long start, unsigned long len, bool enable_wp); +/* move_pages */ +void double_pt_lock(spinlock_t *ptl1, spinlock_t *ptl2); +void double_pt_unlock(spinlock_t *ptl1, spinlock_t *ptl2); +ssize_t move_pages(struct userfaultfd_ctx *ctx, struct mm_struct *mm, + unsigned long dst_start, unsigned long src_start, + unsigned long len, __u64 flags); +int move_pages_huge_pmd(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd, pmd_t dst_pmdval, + struct vm_area_struct *dst_vma, + struct vm_area_struct *src_vma, + unsigned long dst_addr, unsigned long src_addr); + /* mm helpers */ static inline bool is_mergeable_vm_userfaultfd_ctx(struct vm_area_struct *vma, struct vm_userfaultfd_ctx vm_ctx) diff --git a/include/uapi/linux/userfaultfd.h b/include/uapi/linux/userfaultfd.h index 0dbc81015018..2841e4ea8f2c 100644 --- a/include/uapi/linux/userfaultfd.h +++ b/include/uapi/linux/userfaultfd.h @@ -41,7 +41,8 @@ UFFD_FEATURE_WP_HUGETLBFS_SHMEM | \ UFFD_FEATURE_WP_UNPOPULATED | \ UFFD_FEATURE_POISON | \ - UFFD_FEATURE_WP_ASYNC) + UFFD_FEATURE_WP_ASYNC | \ + UFFD_FEATURE_MOVE) #define UFFD_API_IOCTLS \ ((__u64)1 << _UFFDIO_REGISTER | \ (__u64)1 << _UFFDIO_UNREGISTER | \ @@ -50,6 +51,7 @@ ((__u64)1 << _UFFDIO_WAKE | \ (__u64)1 << _UFFDIO_COPY | \ (__u64)1 << _UFFDIO_ZEROPAGE | \ + (__u64)1 << _UFFDIO_MOVE | \ (__u64)1 << _UFFDIO_WRITEPROTECT | \ (__u64)1 << _UFFDIO_CONTINUE | \ (__u64)1 << _UFFDIO_POISON) @@ -73,6 +75,7 @@ #define _UFFDIO_WAKE (0x02) #define _UFFDIO_COPY (0x03) #define _UFFDIO_ZEROPAGE (0x04) +#define _UFFDIO_MOVE (0x05) #define _UFFDIO_WRITEPROTECT (0x06) #define _UFFDIO_CONTINUE (0x07) #define _UFFDIO_POISON (0x08) @@ -92,6 +95,8 @@ struct uffdio_copy) #define UFFDIO_ZEROPAGE _IOWR(UFFDIO, _UFFDIO_ZEROPAGE, \ struct uffdio_zeropage) +#define UFFDIO_MOVE _IOWR(UFFDIO, _UFFDIO_MOVE, \ + struct uffdio_move) #define UFFDIO_WRITEPROTECT _IOWR(UFFDIO, _UFFDIO_WRITEPROTECT, \ struct uffdio_writeprotect) #define UFFDIO_CONTINUE _IOWR(UFFDIO, _UFFDIO_CONTINUE, \ @@ -222,6 +227,9 @@ struct uffdio_api { * asynchronous mode is supported in which the write fault is * automatically resolved and write-protection is un-set. * It implies UFFD_FEATURE_WP_UNPOPULATED. + * + * UFFD_FEATURE_MOVE indicates that the kernel supports moving an + * existing page contents from userspace. */ #define UFFD_FEATURE_PAGEFAULT_FLAG_WP (1<<0) #define UFFD_FEATURE_EVENT_FORK (1<<1) @@ -239,6 +247,7 @@ struct uffdio_api { #define UFFD_FEATURE_WP_UNPOPULATED (1<<13) #define UFFD_FEATURE_POISON (1<<14) #define UFFD_FEATURE_WP_ASYNC (1<<15) +#define UFFD_FEATURE_MOVE (1<<16) __u64 features; __u64 ioctls; @@ -347,6 +356,24 @@ struct uffdio_poison { __s64 updated; }; +struct uffdio_move { + __u64 dst; + __u64 src; + __u64 len; + /* + * Especially if used to atomically remove memory from the + * address space the wake on the dst range is not needed. + */ +#define UFFDIO_MOVE_MODE_DONTWAKE ((__u64)1<<0) +#define UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES ((__u64)1<<1) + __u64 mode; + /* + * "move" is written by the ioctl and must be at the end: the + * copy_from_user will not read the last 8 bytes. + */ + __s64 move; +}; + /* * Flags for the userfaultfd(2) system call itself. */ diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 4f542444a91f..315968db1ca4 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -1964,6 +1964,128 @@ int change_huge_pmd(struct mmu_gather *tlb, struct vm_area_struct *vma, return ret; } +#ifdef CONFIG_USERFAULTFD +/* + * The PT lock for src_pmd and the mmap_lock for reading are held by + * the caller, but it must return after releasing the page_table_lock. + * Just move the page from src_pmd to dst_pmd if possible. + * Return zero if succeeded in moving the page, -EAGAIN if it needs to be + * repeated by the caller, or other errors in case of failure. + */ +int move_pages_huge_pmd(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd, pmd_t dst_pmdval, + struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma, + unsigned long dst_addr, unsigned long src_addr) +{ + pmd_t _dst_pmd, src_pmdval; + struct page *src_page; + struct folio *src_folio; + struct anon_vma *src_anon_vma; + spinlock_t *src_ptl, *dst_ptl; + pgtable_t src_pgtable; + struct mmu_notifier_range range; + int err = 0; + + src_pmdval = *src_pmd; + src_ptl = pmd_lockptr(mm, src_pmd); + + lockdep_assert_held(src_ptl); + mmap_assert_locked(mm); + + /* Sanity checks before the operation */ + if (WARN_ON_ONCE(!pmd_none(dst_pmdval)) || WARN_ON_ONCE(src_addr & ~HPAGE_PMD_MASK) || + WARN_ON_ONCE(dst_addr & ~HPAGE_PMD_MASK)) { + spin_unlock(src_ptl); + return -EINVAL; + } + + if (!pmd_trans_huge(src_pmdval)) { + spin_unlock(src_ptl); + if (is_pmd_migration_entry(src_pmdval)) { + pmd_migration_entry_wait(mm, &src_pmdval); + return -EAGAIN; + } + return -ENOENT; + } + + src_page = pmd_page(src_pmdval); + if (unlikely(!PageAnonExclusive(src_page))) { + spin_unlock(src_ptl); + return -EBUSY; + } + + src_folio = page_folio(src_page); + folio_get(src_folio); + spin_unlock(src_ptl); + + flush_cache_range(src_vma, src_addr, src_addr + HPAGE_PMD_SIZE); + mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, src_addr, + src_addr + HPAGE_PMD_SIZE); + mmu_notifier_invalidate_range_start(&range); + + folio_lock(src_folio); + + /* + * split_huge_page walks the anon_vma chain without the page + * lock. Serialize against it with the anon_vma lock, the page + * lock is not enough. + */ + src_anon_vma = folio_get_anon_vma(src_folio); + if (!src_anon_vma) { + err = -EAGAIN; + goto unlock_folio; + } + anon_vma_lock_write(src_anon_vma); + + dst_ptl = pmd_lockptr(mm, dst_pmd); + double_pt_lock(src_ptl, dst_ptl); + if (unlikely(!pmd_same(*src_pmd, src_pmdval) || + !pmd_same(*dst_pmd, dst_pmdval))) { + err = -EAGAIN; + goto unlock_ptls; + } + if (folio_maybe_dma_pinned(src_folio) || + !PageAnonExclusive(&src_folio->page)) { + err = -EBUSY; + goto unlock_ptls; + } + + if (WARN_ON_ONCE(!folio_test_head(src_folio)) || + WARN_ON_ONCE(!folio_test_anon(src_folio))) { + err = -EBUSY; + goto unlock_ptls; + } + + folio_move_anon_rmap(src_folio, dst_vma); + WRITE_ONCE(src_folio->index, linear_page_index(dst_vma, dst_addr)); + + src_pmdval = pmdp_huge_clear_flush(src_vma, src_addr, src_pmd); + /* Folio got pinned from under us. Put it back and fail the move. */ + if (folio_maybe_dma_pinned(src_folio)) { + set_pmd_at(mm, src_addr, src_pmd, src_pmdval); + err = -EBUSY; + goto unlock_ptls; + } + + _dst_pmd = mk_huge_pmd(&src_folio->page, dst_vma->vm_page_prot); + /* Follow mremap() behavior and treat the entry dirty after the move */ + _dst_pmd = pmd_mkwrite(pmd_mkdirty(_dst_pmd), dst_vma); + set_pmd_at(mm, dst_addr, dst_pmd, _dst_pmd); + + src_pgtable = pgtable_trans_huge_withdraw(mm, src_pmd); + pgtable_trans_huge_deposit(mm, dst_pmd, src_pgtable); +unlock_ptls: + double_pt_unlock(src_ptl, dst_ptl); + anon_vma_unlock_write(src_anon_vma); + put_anon_vma(src_anon_vma); +unlock_folio: + /* unblock rmap walks */ + folio_unlock(src_folio); + mmu_notifier_invalidate_range_end(&range); + folio_put(src_folio); + return err; +} +#endif /* CONFIG_USERFAULTFD */ + /* * Returns page table lock pointer if a given pmd maps a thp, NULL otherwise. * diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 064654717843..0da6937572cf 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -1139,6 +1139,9 @@ static int collapse_huge_page(struct mm_struct *mm, unsigned long address, * Prevent all access to pagetables with the exception of * gup_fast later handled by the ptep_clear_flush and the VM * handled by the anon_vma lock + PG_lock. + * + * UFFDIO_MOVE is prevented to race as well thanks to the + * mmap_lock. */ mmap_write_lock(mm); result = hugepage_vma_revalidate(mm, address, true, &vma, cc); diff --git a/mm/rmap.c b/mm/rmap.c index 525c5bc0b0b3..de9426ad0f1b 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -490,6 +490,12 @@ void __init anon_vma_init(void) * page_remove_rmap() that the anon_vma pointer from page->mapping is valid * if there is a mapcount, we can dereference the anon_vma after observing * those. + * + * NOTE: the caller should normally hold folio lock when calling this. If + * not, the caller needs to double check the anon_vma didn't change after + * taking the anon_vma lock for either read or write (UFFDIO_MOVE can modify it + * concurrently without folio lock protection). See folio_lock_anon_vma_read() + * which has already covered that, and comment above remap_pages(). */ struct anon_vma *folio_get_anon_vma(struct folio *folio) { diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index 0b6ca553bebe..71d0281f1162 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -842,3 +842,602 @@ int mwriteprotect_range(struct mm_struct *dst_mm, unsigned long start, mmap_read_unlock(dst_mm); return err; } + + +void double_pt_lock(spinlock_t *ptl1, + spinlock_t *ptl2) + __acquires(ptl1) + __acquires(ptl2) +{ + spinlock_t *ptl_tmp; + + if (ptl1 > ptl2) { + /* exchange ptl1 and ptl2 */ + ptl_tmp = ptl1; + ptl1 = ptl2; + ptl2 = ptl_tmp; + } + /* lock in virtual address order to avoid lock inversion */ + spin_lock(ptl1); + if (ptl1 != ptl2) + spin_lock_nested(ptl2, SINGLE_DEPTH_NESTING); + else + __acquire(ptl2); +} + +void double_pt_unlock(spinlock_t *ptl1, + spinlock_t *ptl2) + __releases(ptl1) + __releases(ptl2) +{ + spin_unlock(ptl1); + if (ptl1 != ptl2) + spin_unlock(ptl2); + else + __release(ptl2); +} + + +static int move_present_pte(struct mm_struct *mm, + struct vm_area_struct *dst_vma, + struct vm_area_struct *src_vma, + unsigned long dst_addr, unsigned long src_addr, + pte_t *dst_pte, pte_t *src_pte, + pte_t orig_dst_pte, pte_t orig_src_pte, + spinlock_t *dst_ptl, spinlock_t *src_ptl, + struct folio *src_folio) +{ + int err = 0; + + double_pt_lock(dst_ptl, src_ptl); + + if (!pte_same(*src_pte, orig_src_pte) || + !pte_same(*dst_pte, orig_dst_pte)) { + err = -EAGAIN; + goto out; + } + if (folio_test_large(src_folio) || + folio_maybe_dma_pinned(src_folio) || + !PageAnonExclusive(&src_folio->page)) { + err = -EBUSY; + goto out; + } + + folio_move_anon_rmap(src_folio, dst_vma); + WRITE_ONCE(src_folio->index, linear_page_index(dst_vma, dst_addr)); + + orig_src_pte = ptep_clear_flush(src_vma, src_addr, src_pte); + /* Folio got pinned from under us. Put it back and fail the move. */ + if (folio_maybe_dma_pinned(src_folio)) { + set_pte_at(mm, src_addr, src_pte, orig_src_pte); + err = -EBUSY; + goto out; + } + + orig_dst_pte = mk_pte(&src_folio->page, dst_vma->vm_page_prot); + /* Follow mremap() behavior and treat the entry dirty after the move */ + orig_dst_pte = pte_mkwrite(pte_mkdirty(orig_dst_pte), dst_vma); + + set_pte_at(mm, dst_addr, dst_pte, orig_dst_pte); +out: + double_pt_unlock(dst_ptl, src_ptl); + return err; +} + +static int move_swap_pte(struct mm_struct *mm, + unsigned long dst_addr, unsigned long src_addr, + pte_t *dst_pte, pte_t *src_pte, + pte_t orig_dst_pte, pte_t orig_src_pte, + spinlock_t *dst_ptl, spinlock_t *src_ptl) +{ + if (!pte_swp_exclusive(orig_src_pte)) + return -EBUSY; + + double_pt_lock(dst_ptl, src_ptl); + + if (!pte_same(*src_pte, orig_src_pte) || + !pte_same(*dst_pte, orig_dst_pte)) { + double_pt_unlock(dst_ptl, src_ptl); + return -EAGAIN; + } + + orig_src_pte = ptep_get_and_clear(mm, src_addr, src_pte); + set_pte_at(mm, dst_addr, dst_pte, orig_src_pte); + double_pt_unlock(dst_ptl, src_ptl); + + return 0; +} + +/* + * The mmap_lock for reading is held by the caller. Just move the page + * from src_pmd to dst_pmd if possible, and return true if succeeded + * in moving the page. + */ +static int move_pages_pte(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd, + struct vm_area_struct *dst_vma, + struct vm_area_struct *src_vma, + unsigned long dst_addr, unsigned long src_addr, + __u64 mode) +{ + swp_entry_t entry; + pte_t orig_src_pte, orig_dst_pte; + pte_t src_folio_pte; + spinlock_t *src_ptl, *dst_ptl; + pte_t *src_pte = NULL; + pte_t *dst_pte = NULL; + + struct folio *src_folio = NULL; + struct anon_vma *src_anon_vma = NULL; + struct mmu_notifier_range range; + int err = 0; + + flush_cache_range(src_vma, src_addr, src_addr + PAGE_SIZE); + mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, mm, + src_addr, src_addr + PAGE_SIZE); + mmu_notifier_invalidate_range_start(&range); +retry: + dst_pte = pte_offset_map_nolock(mm, dst_pmd, dst_addr, &dst_ptl); + + /* Retry if a huge pmd materialized from under us */ + if (unlikely(!dst_pte)) { + err = -EAGAIN; + goto out; + } + + src_pte = pte_offset_map_nolock(mm, src_pmd, src_addr, &src_ptl); + + /* + * We held the mmap_lock for reading so MADV_DONTNEED + * can zap transparent huge pages under us, or the + * transparent huge page fault can establish new + * transparent huge pages under us. + */ + if (unlikely(!src_pte)) { + err = -EAGAIN; + goto out; + } + + /* Sanity checks before the operation */ + if (WARN_ON_ONCE(pmd_none(*dst_pmd)) || WARN_ON_ONCE(pmd_none(*src_pmd)) || + WARN_ON_ONCE(pmd_trans_huge(*dst_pmd)) || WARN_ON_ONCE(pmd_trans_huge(*src_pmd))) { + err = -EINVAL; + goto out; + } + + spin_lock(dst_ptl); + orig_dst_pte = *dst_pte; + spin_unlock(dst_ptl); + if (!pte_none(orig_dst_pte)) { + err = -EEXIST; + goto out; + } + + spin_lock(src_ptl); + orig_src_pte = *src_pte; + spin_unlock(src_ptl); + if (pte_none(orig_src_pte)) { + if (!(mode & UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES)) + err = -ENOENT; + else /* nothing to do to move a hole */ + err = 0; + goto out; + } + + /* If PTE changed after we locked the folio them start over */ + if (src_folio && unlikely(!pte_same(src_folio_pte, orig_src_pte))) { + err = -EAGAIN; + goto out; + } + + if (pte_present(orig_src_pte)) { + /* + * Pin and lock both source folio and anon_vma. Since we are in + * RCU read section, we can't block, so on contention have to + * unmap the ptes, obtain the lock and retry. + */ + if (!src_folio) { + struct folio *folio; + + /* + * Pin the page while holding the lock to be sure the + * page isn't freed under us + */ + spin_lock(src_ptl); + if (!pte_same(orig_src_pte, *src_pte)) { + spin_unlock(src_ptl); + err = -EAGAIN; + goto out; + } + + folio = vm_normal_folio(src_vma, src_addr, orig_src_pte); + if (!folio || folio_test_large(folio) || + !PageAnonExclusive(&folio->page)) { + spin_unlock(src_ptl); + err = -EBUSY; + goto out; + } + + folio_get(folio); + src_folio = folio; + src_folio_pte = orig_src_pte; + spin_unlock(src_ptl); + + if (!folio_trylock(src_folio)) { + pte_unmap(&orig_src_pte); + pte_unmap(&orig_dst_pte); + src_pte = dst_pte = NULL; + /* now we can block and wait */ + folio_lock(src_folio); + goto retry; + } + + if (WARN_ON_ONCE(!folio_test_anon(src_folio))) { + err = -EBUSY; + goto out; + } + } + + if (!src_anon_vma) { + /* + * folio_referenced walks the anon_vma chain + * without the folio lock. Serialize against it with + * the anon_vma lock, the folio lock is not enough. + */ + src_anon_vma = folio_get_anon_vma(src_folio); + if (!src_anon_vma) { + /* page was unmapped from under us */ + err = -EAGAIN; + goto out; + } + if (!anon_vma_trylock_write(src_anon_vma)) { + pte_unmap(&orig_src_pte); + pte_unmap(&orig_dst_pte); + src_pte = dst_pte = NULL; + /* now we can block and wait */ + anon_vma_lock_write(src_anon_vma); + goto retry; + } + } + + err = move_present_pte(mm, dst_vma, src_vma, + dst_addr, src_addr, dst_pte, src_pte, + orig_dst_pte, orig_src_pte, + dst_ptl, src_ptl, src_folio); + } else { + entry = pte_to_swp_entry(orig_src_pte); + if (non_swap_entry(entry)) { + if (is_migration_entry(entry)) { + pte_unmap(&orig_src_pte); + pte_unmap(&orig_dst_pte); + src_pte = dst_pte = NULL; + migration_entry_wait(mm, src_pmd, src_addr); + err = -EAGAIN; + } else + err = -EFAULT; + goto out; + } + + err = move_swap_pte(mm, dst_addr, src_addr, + dst_pte, src_pte, + orig_dst_pte, orig_src_pte, + dst_ptl, src_ptl); + } + +out: + if (src_anon_vma) { + anon_vma_unlock_write(src_anon_vma); + put_anon_vma(src_anon_vma); + } + if (src_folio) { + folio_unlock(src_folio); + folio_put(src_folio); + } + if (dst_pte) + pte_unmap(dst_pte); + if (src_pte) + pte_unmap(src_pte); + mmu_notifier_invalidate_range_end(&range); + + return err; +} + +#ifdef CONFIG_TRANSPARENT_HUGEPAGE +static inline bool move_splits_huge_pmd(unsigned long dst_addr, + unsigned long src_addr, + unsigned long src_end) +{ + return (src_addr & ~HPAGE_PMD_MASK) || (dst_addr & ~HPAGE_PMD_MASK) || + src_end - src_addr < HPAGE_PMD_SIZE; +} +#else +static inline bool move_splits_huge_pmd(unsigned long dst_addr, + unsigned long src_addr, + unsigned long src_end) +{ + /* This is unreachable anyway, just to avoid warnings when HPAGE_PMD_SIZE==0 */ + return false; +} +#endif + +static inline bool vma_move_compatible(struct vm_area_struct *vma) +{ + return !(vma->vm_flags & (VM_PFNMAP | VM_IO | VM_HUGETLB | + VM_MIXEDMAP | VM_SHADOW_STACK)); +} + +static int validate_move_areas(struct userfaultfd_ctx *ctx, + struct vm_area_struct *src_vma, + struct vm_area_struct *dst_vma) +{ + /* Only allow moving if both have the same access and protection */ + if ((src_vma->vm_flags & VM_ACCESS_FLAGS) != (dst_vma->vm_flags & VM_ACCESS_FLAGS) || + pgprot_val(src_vma->vm_page_prot) != pgprot_val(dst_vma->vm_page_prot)) + return -EINVAL; + + /* Only allow moving if both are mlocked or both aren't */ + if ((src_vma->vm_flags & VM_LOCKED) != (dst_vma->vm_flags & VM_LOCKED)) + return -EINVAL; + + /* + * For now, we keep it simple and only move between writable VMAs. + * Access flags are equal, therefore cheching only the source is enough. + */ + if (!(src_vma->vm_flags & VM_WRITE)) + return -EINVAL; + + /* Check if vma flags indicate content which can be moved */ + if (!vma_move_compatible(src_vma) || !vma_move_compatible(dst_vma)) + return -EINVAL; + + /* Ensure dst_vma is registered in uffd we are operating on */ + if (!dst_vma->vm_userfaultfd_ctx.ctx || + dst_vma->vm_userfaultfd_ctx.ctx != ctx) + return -EINVAL; + + /* Only allow moving across anonymous vmas */ + if (!vma_is_anonymous(src_vma) || !vma_is_anonymous(dst_vma)) + return -EINVAL; + + /* + * Ensure the dst_vma has a anon_vma or this page + * would get a NULL anon_vma when moved in the + * dst_vma. + */ + if (unlikely(anon_vma_prepare(dst_vma))) + return -ENOMEM; + + return 0; +} + +/** + * move_pages - move arbitrary anonymous pages of an existing vma + * @ctx: pointer to the userfaultfd context + * @mm: the address space to move pages + * @dst_start: start of the destination virtual memory range + * @src_start: start of the source virtual memory range + * @len: length of the virtual memory range + * @mode: flags from uffdio_move.mode + * + * Must be called with mmap_lock held for read. + * + * move_pages() remaps arbitrary anonymous pages atomically in zero + * copy. It only works on non shared anonymous pages because those can + * be relocated without generating non linear anon_vmas in the rmap + * code. + * + * It provides a zero copy mechanism to handle userspace page faults. + * The source vma pages should have mapcount == 1, which can be + * enforced by using madvise(MADV_DONTFORK) on src vma. + * + * The thread receiving the page during the userland page fault + * will receive the faulting page in the source vma through the network, + * storage or any other I/O device (MADV_DONTFORK in the source vma + * avoids move_pages() to fail with -EBUSY if the process forks before + * move_pages() is called), then it will call move_pages() to map the + * page in the faulting address in the destination vma. + * + * This userfaultfd command works purely via pagetables, so it's the + * most efficient way to move physical non shared anonymous pages + * across different virtual addresses. Unlike mremap()/mmap()/munmap() + * it does not create any new vmas. The mapping in the destination + * address is atomic. + * + * It only works if the vma protection bits are identical from the + * source and destination vma. + * + * It can remap non shared anonymous pages within the same vma too. + * + * If the source virtual memory range has any unmapped holes, or if + * the destination virtual memory range is not a whole unmapped hole, + * move_pages() will fail respectively with -ENOENT or -EEXIST. This + * provides a very strict behavior to avoid any chance of memory + * corruption going unnoticed if there are userland race conditions. + * Only one thread should resolve the userland page fault at any given + * time for any given faulting address. This means that if two threads + * try to both call move_pages() on the same destination address at the + * same time, the second thread will get an explicit error from this + * command. + * + * The command retval will return "len" is successful. The command + * however can be interrupted by fatal signals or errors. If + * interrupted it will return the number of bytes successfully + * remapped before the interruption if any, or the negative error if + * none. It will never return zero. Either it will return an error or + * an amount of bytes successfully moved. If the retval reports a + * "short" remap, the move_pages() command should be repeated by + * userland with src+retval, dst+reval, len-retval if it wants to know + * about the error that interrupted it. + * + * The UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES flag can be specified to + * prevent -ENOENT errors to materialize if there are holes in the + * source virtual range that is being remapped. The holes will be + * accounted as successfully remapped in the retval of the + * command. This is mostly useful to remap hugepage naturally aligned + * virtual regions without knowing if there are transparent hugepage + * in the regions or not, but preventing the risk of having to split + * the hugepmd during the remap. + * + * If there's any rmap walk that is taking the anon_vma locks without + * first obtaining the folio lock (the only current instance is + * folio_referenced), they will have to verify if the folio->mapping + * has changed after taking the anon_vma lock. If it changed they + * should release the lock and retry obtaining a new anon_vma, because + * it means the anon_vma was changed by move_pages() before the lock + * could be obtained. This is the only additional complexity added to + * the rmap code to provide this anonymous page remapping functionality. + */ +ssize_t move_pages(struct userfaultfd_ctx *ctx, struct mm_struct *mm, + unsigned long dst_start, unsigned long src_start, + unsigned long len, __u64 mode) +{ + struct vm_area_struct *src_vma, *dst_vma; + unsigned long src_addr, dst_addr; + pmd_t *src_pmd, *dst_pmd; + long err = -EINVAL; + ssize_t moved = 0; + + /* Sanitize the command parameters. */ + if (WARN_ON_ONCE(src_start & ~PAGE_MASK) || + WARN_ON_ONCE(dst_start & ~PAGE_MASK) || + WARN_ON_ONCE(len & ~PAGE_MASK)) + goto out; + + /* Does the address range wrap, or is the span zero-sized? */ + if (WARN_ON_ONCE(src_start + len <= src_start) || + WARN_ON_ONCE(dst_start + len <= dst_start)) + goto out; + + /* + * Make sure the vma is not shared, that the src and dst remap + * ranges are both valid and fully within a single existing + * vma. + */ + src_vma = find_vma(mm, src_start); + if (!src_vma || (src_vma->vm_flags & VM_SHARED)) + goto out; + if (src_start < src_vma->vm_start || + src_start + len > src_vma->vm_end) + goto out; + + dst_vma = find_vma(mm, dst_start); + if (!dst_vma || (dst_vma->vm_flags & VM_SHARED)) + goto out; + if (dst_start < dst_vma->vm_start || + dst_start + len > dst_vma->vm_end) + goto out; + + err = validate_move_areas(ctx, src_vma, dst_vma); + if (err) + goto out; + + for (src_addr = src_start, dst_addr = dst_start; + src_addr < src_start + len;) { + spinlock_t *ptl; + pmd_t dst_pmdval; + unsigned long step_size; + + /* + * Below works because anonymous area would not have a + * transparent huge PUD. If file-backed support is added, + * that case would need to be handled here. + */ + src_pmd = mm_find_pmd(mm, src_addr); + if (unlikely(!src_pmd)) { + if (!(mode & UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES)) { + err = -ENOENT; + break; + } + src_pmd = mm_alloc_pmd(mm, src_addr); + if (unlikely(!src_pmd)) { + err = -ENOMEM; + break; + } + } + dst_pmd = mm_alloc_pmd(mm, dst_addr); + if (unlikely(!dst_pmd)) { + err = -ENOMEM; + break; + } + + dst_pmdval = pmdp_get_lockless(dst_pmd); + /* + * If the dst_pmd is mapped as THP don't override it and just + * be strict. If dst_pmd changes into TPH after this check, the + * move_pages_huge_pmd() will detect the change and retry + * while move_pages_pte() will detect the change and fail. + */ + if (unlikely(pmd_trans_huge(dst_pmdval))) { + err = -EEXIST; + break; + } + + ptl = pmd_trans_huge_lock(src_pmd, src_vma); + if (ptl) { + if (pmd_devmap(*src_pmd)) { + spin_unlock(ptl); + err = -ENOENT; + break; + } + + /* Check if we can move the pmd without splitting it. */ + if (move_splits_huge_pmd(dst_addr, src_addr, src_start + len) || + !pmd_none(dst_pmdval)) { + spin_unlock(ptl); + split_huge_pmd(src_vma, src_pmd, src_addr); + continue; + } + + err = move_pages_huge_pmd(mm, dst_pmd, src_pmd, + dst_pmdval, dst_vma, src_vma, + dst_addr, src_addr); + step_size = HPAGE_PMD_SIZE; + } else { + if (pmd_none(*src_pmd)) { + if (!(mode & UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES)) { + err = -ENOENT; + break; + } + if (unlikely(__pte_alloc(mm, src_pmd))) { + err = -ENOMEM; + break; + } + } + + if (unlikely(pte_alloc(mm, dst_pmd))) { + err = -ENOMEM; + break; + } + + err = move_pages_pte(mm, dst_pmd, src_pmd, + dst_vma, src_vma, + dst_addr, src_addr, mode); + step_size = PAGE_SIZE; + } + + cond_resched(); + + if (fatal_signal_pending(current)) { + /* Do not override an error */ + if (!err || err == -EAGAIN) + err = -EINTR; + break; + } + + if (err) { + if (err == -EAGAIN) + continue; + break; + } + + /* Proceed to the next page */ + dst_addr += step_size; + src_addr += step_size; + moved += step_size; + } + +out: + VM_WARN_ON(moved < 0); + VM_WARN_ON(err > 0); + VM_WARN_ON(!moved && !err); + return moved ? moved : err; +} From patchwork Tue Nov 21 17:16:36 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Suren Baghdasaryan X-Patchwork-Id: 13463395 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1785FC61D85 for ; Tue, 21 Nov 2023 17:16:58 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 580C06B04C6; Tue, 21 Nov 2023 12:16:57 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id 50A3F6B04C9; Tue, 21 Nov 2023 12:16:57 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 35CB46B04CB; Tue, 21 Nov 2023 12:16:57 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0014.hostedemail.com [216.40.44.14]) by kanga.kvack.org (Postfix) with ESMTP id 218266B04C6 for ; Tue, 21 Nov 2023 12:16:57 -0500 (EST) Received: from smtpin02.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay05.hostedemail.com (Postfix) with ESMTP id EE8CE40A82 for ; Tue, 21 Nov 2023 17:16:56 +0000 (UTC) X-FDA: 81482616432.02.D8E5F3C Received: from mail-yw1-f202.google.com (mail-yw1-f202.google.com [209.85.128.202]) by imf28.hostedemail.com (Postfix) with ESMTP id 33458C0008 for ; Tue, 21 Nov 2023 17:16:54 +0000 (UTC) Authentication-Results: imf28.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b="tx7CD/B2"; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf28.hostedemail.com: domain of 3BuZcZQYKCHcnpmZiWbjjbgZ.Xjhgdips-hhfqVXf.jmb@flex--surenb.bounces.google.com designates 209.85.128.202 as permitted sender) smtp.mailfrom=3BuZcZQYKCHcnpmZiWbjjbgZ.Xjhgdips-hhfqVXf.jmb@flex--surenb.bounces.google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1700587015; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=+oQWC9tOLeE7lomKytBDg6Tp5S7jbgII3AVJArGJfwY=; b=VgWKsG+gyGchBg3fgvdFG/yH1+M1Up74CTr04ue4mBje2sRxcYHMfl3rdKT9i6z+N3ea6B pmBEUG7IEd8rTHvGyq7mcFLneneeW0a3ncV3TvzKPQFzFbmKKrjE6GxLDNQctV7/cBwyj8 jafv9fm8eel3R7QvHxdH7GEobYhIlmI= ARC-Authentication-Results: i=1; imf28.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b="tx7CD/B2"; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf28.hostedemail.com: domain of 3BuZcZQYKCHcnpmZiWbjjbgZ.Xjhgdips-hhfqVXf.jmb@flex--surenb.bounces.google.com designates 209.85.128.202 as permitted sender) smtp.mailfrom=3BuZcZQYKCHcnpmZiWbjjbgZ.Xjhgdips-hhfqVXf.jmb@flex--surenb.bounces.google.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700587015; a=rsa-sha256; cv=none; b=fgfnoFIX5ZIhgBJQvCwDmQ6x243G9dfsRoOLkZ1pkNbl8hNAB1nI7H5IEUbzChuk1Rw53C eS1LlnvJoJ8Ve9hSNrzb1LQcC+bQ1W9ecLD8iSqPcps4g03yTPMVDTyfY0qWQNRvPTI50R ZH2nrWjYBhv46jb/WbArAnQnhdbfcc4= Received: by mail-yw1-f202.google.com with SMTP id 00721157ae682-5c9e6c37bc4so39523717b3.2 for ; Tue, 21 Nov 2023 09:16:54 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1700587014; x=1701191814; darn=kvack.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=+oQWC9tOLeE7lomKytBDg6Tp5S7jbgII3AVJArGJfwY=; b=tx7CD/B2/8o773w32XW+kqnTve+6TxR57+fh6+i1YoPSEnsX0i1rCoOdIStk0/DF+N uxZFD5grXn0Wi5UShzDTdrPa0iO903xk64ASpSKiUyMPVEvbVEZraca4b7PFh2sUSI0L EcOCpaK3s4DcrT8G26c9BprLrqSgv9drJNED//s/dcE2HVHndp7nNYPAYhuFqEJUJpQV v45ZNUJoKccxG3OfgWBEQqSVVTSVHmqrb7k0DD08R4NbI4Fy1PlklUgVLcuCR3avtlZc dL8O3EVMqU7fVtdTeHOThz1xNsV9+sGBpPriQlXjYqy18DX8nBqxdmAGRUjwpUfh8Gcw VdFQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1700587014; x=1701191814; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=+oQWC9tOLeE7lomKytBDg6Tp5S7jbgII3AVJArGJfwY=; b=Dj24SKqaUWbMSoUzhcqwGSPlU/wbX/qjS55MqN8hjkEyld25jaEnObwz3alQD5OYh+ xWTyfWyNt8bk+Anf/xCKJCAvVXjXFG5uuVje+Mt9S04TJFOiAhisX6kO0SjVMRKoYBg2 vuMSx7qaH+YrE+aVdPbEB3h2xis8+0ijDYMpSKsRY8YvfKdN9C98f02beemkBk9o2Ytj URaBxzRVeqTYXUrEMOPYTqxHzaNBXlusKzcyRUOHDeRDV/I5SPq+WG6fvoV1Snv9vexm 1irf4Y2BQPkRUwU1uTWBu0H0XpwlRGDfJNn2N4BZUVUwDqPRHmMAnVjFZ338ENGN+KMl VXzA== X-Gm-Message-State: AOJu0YzYv+ckgv7FVGyM/0MdW8cAH1ve1MZphMQuvybtwgIQJu2Jr9cq RbSR/XrUs72h3uL4sQ7sar+ToGG1v7s= X-Google-Smtp-Source: AGHT+IGt4XeQzmmYh3UGYq1fZSYeCCXt9DvRPaCqVRLnjG7738p7ZSz3vsLWilwzFfU1DtxScxBgR2AWR/c= X-Received: from surenb-desktop.mtv.corp.google.com ([2620:15c:211:201:2045:f6d2:f01d:3fff]) (user=surenb job=sendgmr) by 2002:a05:690c:891:b0:5c9:1c6a:40f with SMTP id cd17-20020a05690c089100b005c91c6a040fmr290590ywb.5.1700587014155; Tue, 21 Nov 2023 09:16:54 -0800 (PST) Date: Tue, 21 Nov 2023 09:16:36 -0800 In-Reply-To: <20231121171643.3719880-1-surenb@google.com> Mime-Version: 1.0 References: <20231121171643.3719880-1-surenb@google.com> X-Mailer: git-send-email 2.43.0.rc1.413.gea7ed67945-goog Message-ID: <20231121171643.3719880-4-surenb@google.com> Subject: [PATCH v5 3/5] selftests/mm: call uffd_test_ctx_clear at the end of the test From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: viro@zeniv.linux.org.uk, brauner@kernel.org, shuah@kernel.org, aarcange@redhat.com, lokeshgidra@google.com, peterx@redhat.com, david@redhat.com, hughd@google.com, mhocko@suse.com, axelrasmussen@google.com, rppt@kernel.org, willy@infradead.org, Liam.Howlett@oracle.com, jannh@google.com, zhangpeng362@huawei.com, bgeffon@google.com, kaleshsingh@google.com, ngeoffray@google.com, jdduke@google.com, surenb@google.com, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, kernel-team@android.com X-Rspam-User: X-Rspamd-Server: rspam12 X-Rspamd-Queue-Id: 33458C0008 X-Stat-Signature: r95ntx6ajcbatyay4ahawjjuqfoxscni X-HE-Tag: 1700587014-563189 X-HE-Meta: U2FsdGVkX19TQSVqjpz3c4WWWfynT8EPp4iesEzF7R76a+xNpnsTUoSlMHcszTEC53jqYK5g4iZ7UiBLLr5cV2L8XNHkZvwvPJe1TKGh33e3kyIBNwiQ7U1C4pyO4R4QxW37qSfTWpfkIOV9FlvybhGN9NTSA7X92hpxBryiUtfN/HQVLzIeMNF/dSKjn5PEvTRkQrItezGzuig3+ibRtjkNjry1ixqWtJSH0GkPYVJuO2CJazTiKs7osjRoJySfgjYBRjNY/jpPRY7RkmyqMop4S9uNBOxizEYmyRfiVbzPq89IEaD5TXXEqMaBYk5JFLvT2sarIw3M4I3Y0qqSrxZI5p2RZHwdf8Rb/Q7O6f95dtbbYtlnduaP2er56ZlEK6nJ/FPSNT6wsNFIY/JG/jvKEBit/pZWQK2Gj3dG2DeHYQOw75ssTGd+Jf6Ayth6yr17m5QGySFF2uG1FdcJAi+eV78+1NRnKRZTu+gUm7Aini8IoOfXGbPE168p/mibKV2Ec7oKzqjHvZoQCB2rIAnohvsgjKJQGM5+Qbzd5ox+eDoOHElAVNOo0A1EdMbmgVps5X6oEP03Lczg7LIVht14JcgPxi5Pn2kMTZI3GR8+4I+nkgrAmh6OQPzLNhYTw1TeLoSvEIKqN2q9tbPeBIT/15/L7dS1EdAwficLKYQE7DzXMX/c/FC+hPmUiUv8GPzHAiaB373SxUicVpMHoi+MVX/5d9/BqPTuFrTaCSLHDWL1kzYKJEKcIPl252myM/wxUXKg1sGf/Lwg5cHY7Iu8JgkhjPRgVlt1DLEs5QNtZS62Wu0YCaaGW0+Z8rKFCODDpZkI122YrmeMNv+pfi+nWW/c8YmAKaIF3GrY93iba+hJCuYob/6+3vips0oNoMLklRiX9spvO1ld4hrBGJwdgQeeN1iQUpBuxQ2oq3ksgH37vt3cgBlJ/2C2TO8PIKjaJcVc8wwVa6eHro1 u7dJGdAB Mv7sMsyJME7+v0mnCy9wfhqEjJlpkvx8OCxIvqZqqDhxgVkVGCEGC5mzUyqbvTuf4ak3CszztDO2P+ku3sYaGBpq1nZq0G6dCvmktOw9+DMpDzwsNchpm7XEe+5vkWhqlbHF30ir1I5KRwD5sgg0/Hh7Mgg+g20crsCqTSWkUOpGjql75ZYWG/zYTuRMO4cUvr0RI3ewwecYHcvarI+qu8DKN1sr9sQEhNYfWbh66VIngBDw0DCCrNlql4Hh5I4staxgoGaBhMvk1yy8BYWyqtvQD4aPOEyHjJJECTbYVUE5S/3FfT/xRdCkT/VInoGbh7ehNEWWNR7QZ8Y0/s2rHgksgYh2pz4a+aXudSLPeGfZFbyqIQEVhDrlPcNRao18nq+uiN7UEQXkvueDgTVxIKJzSZyccG3FEMMPVJ2bCJbT+mL4QPPn7dQ8wynDB8u/cyE2xzPNLmeZ5hwiFFOILSy9slzkvx3fMBRJOd0zJKbmk86LzHDTXDnDfC4d6SCniHVvMOLXfAX9YgUJH/YfR8skKzFnYTNX+ruanOTBQdqlupAsB/9FBu8mFj0s96/odOXQofzQWBMfBaZHnI7sdtCy3NQ5bmMHWV/LFN/101O63uDG1jk2gMzBgoRvlG/VQQE2Vsc8L5N8GSgH1b68ZhhCAIb1EvS3k9usF61MZW+ArmV+6Lq1HzCzDt2TrTLQkLawVc8L5RR465STQaHxSmoqn1A== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: uffd_test_ctx_clear() is being called from uffd_test_ctx_init() to unmap areas used in the previous test run. This approach is problematic because while unmapping areas uffd_test_ctx_clear() uses page_size and nr_pages which might differ from one test run to another. Fix this by calling uffd_test_ctx_clear() after each test is done. Signed-off-by: Suren Baghdasaryan Reviewed-by: Peter Xu Reviewed-by: Axel Rasmussen --- tools/testing/selftests/mm/uffd-common.c | 4 +--- tools/testing/selftests/mm/uffd-common.h | 1 + tools/testing/selftests/mm/uffd-stress.c | 5 ++++- tools/testing/selftests/mm/uffd-unit-tests.c | 1 + 4 files changed, 7 insertions(+), 4 deletions(-) diff --git a/tools/testing/selftests/mm/uffd-common.c b/tools/testing/selftests/mm/uffd-common.c index 02b89860e193..583e5a4cc0fd 100644 --- a/tools/testing/selftests/mm/uffd-common.c +++ b/tools/testing/selftests/mm/uffd-common.c @@ -262,7 +262,7 @@ static inline void munmap_area(void **area) *area = NULL; } -static void uffd_test_ctx_clear(void) +void uffd_test_ctx_clear(void) { size_t i; @@ -298,8 +298,6 @@ int uffd_test_ctx_init(uint64_t features, const char **errmsg) unsigned long nr, cpu; int ret; - uffd_test_ctx_clear(); - ret = uffd_test_ops->allocate_area((void **)&area_src, true); ret |= uffd_test_ops->allocate_area((void **)&area_dst, false); if (ret) { diff --git a/tools/testing/selftests/mm/uffd-common.h b/tools/testing/selftests/mm/uffd-common.h index 7c4fa964c3b0..870776b5a323 100644 --- a/tools/testing/selftests/mm/uffd-common.h +++ b/tools/testing/selftests/mm/uffd-common.h @@ -105,6 +105,7 @@ extern uffd_test_ops_t *uffd_test_ops; void uffd_stats_report(struct uffd_args *args, int n_cpus); int uffd_test_ctx_init(uint64_t features, const char **errmsg); +void uffd_test_ctx_clear(void); int userfaultfd_open(uint64_t *features); int uffd_read_msg(int ufd, struct uffd_msg *msg); void wp_range(int ufd, __u64 start, __u64 len, bool wp); diff --git a/tools/testing/selftests/mm/uffd-stress.c b/tools/testing/selftests/mm/uffd-stress.c index 469e0476af26..7e83829bbb33 100644 --- a/tools/testing/selftests/mm/uffd-stress.c +++ b/tools/testing/selftests/mm/uffd-stress.c @@ -323,8 +323,10 @@ static int userfaultfd_stress(void) uffd_stats_reset(args, nr_cpus); /* bounce pass */ - if (stress(args)) + if (stress(args)) { + uffd_test_ctx_clear(); return 1; + } /* Clear all the write protections if there is any */ if (test_uffdio_wp) @@ -354,6 +356,7 @@ static int userfaultfd_stress(void) uffd_stats_report(args, nr_cpus); } + uffd_test_ctx_clear(); return 0; } diff --git a/tools/testing/selftests/mm/uffd-unit-tests.c b/tools/testing/selftests/mm/uffd-unit-tests.c index 2709a34a39c5..e7d43c198041 100644 --- a/tools/testing/selftests/mm/uffd-unit-tests.c +++ b/tools/testing/selftests/mm/uffd-unit-tests.c @@ -1319,6 +1319,7 @@ int main(int argc, char *argv[]) continue; } test->uffd_fn(&args); + uffd_test_ctx_clear(); } } From patchwork Tue Nov 21 17:16:37 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Suren Baghdasaryan X-Patchwork-Id: 13463396 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id D0E50C61D90 for ; Tue, 21 Nov 2023 17:17:00 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id A63EE6B04CB; Tue, 21 Nov 2023 12:16:59 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id A13916B04CC; Tue, 21 Nov 2023 12:16:59 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 88DD96B04CD; Tue, 21 Nov 2023 12:16:59 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0016.hostedemail.com [216.40.44.16]) by kanga.kvack.org (Postfix) with ESMTP id 6DBF86B04CB for ; Tue, 21 Nov 2023 12:16:59 -0500 (EST) Received: from smtpin02.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay09.hostedemail.com (Postfix) with ESMTP id 4A5D580A8E for ; Tue, 21 Nov 2023 17:16:59 +0000 (UTC) X-FDA: 81482616558.02.170B74A Received: from mail-yw1-f202.google.com (mail-yw1-f202.google.com [209.85.128.202]) by imf16.hostedemail.com (Postfix) with ESMTP id 497DB180009 for ; Tue, 21 Nov 2023 17:16:57 +0000 (UTC) Authentication-Results: imf16.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=RmGgNM8w; spf=pass (imf16.hostedemail.com: domain of 3COZcZQYKCHkprobkYdlldib.Zljifkru-jjhsXZh.lod@flex--surenb.bounces.google.com designates 209.85.128.202 as permitted sender) smtp.mailfrom=3COZcZQYKCHkprobkYdlldib.Zljifkru-jjhsXZh.lod@flex--surenb.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1700587017; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=16WWsVx2wninDVgD360hRJytZJiRJbLfp5hnTHjwaT4=; b=j3t0CHVgzTw1ZEngLpaT/TVEBWiSfP0OsvRJs9IBqwQ/P3phdRAlnzaCpI/nHpWyB7JpDD qJGNS6JuqKpmP2+ofHcrezIXfUajAWeIf9iFTCoeTcVEuES66ziq8fRqwx+21UF301JNym 4wilVcHTe2bxX10zRVhm90xkqBIAiag= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700587017; a=rsa-sha256; cv=none; b=so1D4vFoQd8HtqPTTxTiXxA4PXws3iFDB/FvK5NapJ8+f8PucgmT2i2J8pp87h11LDXv9z vk43uTVOQXCHdMnIe5GyncnDnBwEjIjOIqX4GooepoSqkQ4mWU3NbhrY0y/RRTh3YRpprc tGalfkg8bi6N7h2watvyzQ/Sc4xlUQA= ARC-Authentication-Results: i=1; imf16.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=RmGgNM8w; spf=pass (imf16.hostedemail.com: domain of 3COZcZQYKCHkprobkYdlldib.Zljifkru-jjhsXZh.lod@flex--surenb.bounces.google.com designates 209.85.128.202 as permitted sender) smtp.mailfrom=3COZcZQYKCHkprobkYdlldib.Zljifkru-jjhsXZh.lod@flex--surenb.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com Received: by mail-yw1-f202.google.com with SMTP id 00721157ae682-5ca713d53f3so45739987b3.1 for ; Tue, 21 Nov 2023 09:16:57 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1700587016; x=1701191816; darn=kvack.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=16WWsVx2wninDVgD360hRJytZJiRJbLfp5hnTHjwaT4=; b=RmGgNM8wlicvAoYi5XSOIZk1M7v8m7s+xcYWAWbfvCjsklseDfNgw8Y7P2k3fGKoKE w5P2Ky7Z6vRjJyWgoaLQ6M16QriMpdJPZyc+rxEzyQUATJWnaUyO5pvTKGoF/OavCnIf 8lcxeNcFeAhLIzeEChI8gGJTf3pN2Navjhj8lZALCVZKWlPUgpGPRAOttS2hM4tBWpRV FV5W6973mDrjU6eyss82TuYh8mqyI2YB2K52cJCoiyoT1sqcwC86jeEmELmzDfUcmwkA 7TBM8TtsG8TLYXcjh6gPXfSSGtsWvbiRa/4GrVFEuyQ62nJYprZ8Y8sGxabh6ca5zqy2 b6+Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1700587016; x=1701191816; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=16WWsVx2wninDVgD360hRJytZJiRJbLfp5hnTHjwaT4=; b=BNKx4p1BUMyhUTSXLbA+e1abvFpmE6uXBWtebVMZMu9ANFsAyTwn1jIfn1BvmRXPom vFpDy8R1ER2Zzbq36UwOf8PuOYVqs7zEhqW7jni78bj3Rs0oKWjYbC4B1R4Gc7o7XqRc DlhYZCVAWsRFtyrdcxOHgcJpgsWeA+1y3lJ0qn/lz/intAEYKJh10RxYhjzUwqOtKwcc EkzL1py78sVMCcq44w/iqWV62G83RxYwX4kHRTekup4Bf2qWwdX5JqyHK03IEFsi7Kxv RAL/GS8878KHLN+uzqgnnMn0BGogi8s26iaCDfqQC5WWM1+piThbfHiazOrn+rDtmVm5 oINw== X-Gm-Message-State: AOJu0Yyo+6Z4terRw4OE9E2DFijto3ruA4tR9h80RDMqYWoLqt4yopGZ nlMFbfNFmM9WlRxK8dX1JSz5V9eCqKE= X-Google-Smtp-Source: AGHT+IEDzZp+znMIvLmLUboaHGb6wcOjpm4sLpEi36+gWsJpIjPPrJpEev3RGtNgfOw04VbiAOuwWaVIAJQ= X-Received: from surenb-desktop.mtv.corp.google.com ([2620:15c:211:201:2045:f6d2:f01d:3fff]) (user=surenb job=sendgmr) by 2002:a05:6902:110:b0:d9a:cbf9:1c8d with SMTP id o16-20020a056902011000b00d9acbf91c8dmr324632ybh.12.1700587016415; Tue, 21 Nov 2023 09:16:56 -0800 (PST) Date: Tue, 21 Nov 2023 09:16:37 -0800 In-Reply-To: <20231121171643.3719880-1-surenb@google.com> Mime-Version: 1.0 References: <20231121171643.3719880-1-surenb@google.com> X-Mailer: git-send-email 2.43.0.rc1.413.gea7ed67945-goog Message-ID: <20231121171643.3719880-5-surenb@google.com> Subject: [PATCH v5 4/5] selftests/mm: add uffd_test_case_ops to allow test case-specific operations From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: viro@zeniv.linux.org.uk, brauner@kernel.org, shuah@kernel.org, aarcange@redhat.com, lokeshgidra@google.com, peterx@redhat.com, david@redhat.com, hughd@google.com, mhocko@suse.com, axelrasmussen@google.com, rppt@kernel.org, willy@infradead.org, Liam.Howlett@oracle.com, jannh@google.com, zhangpeng362@huawei.com, bgeffon@google.com, kaleshsingh@google.com, ngeoffray@google.com, jdduke@google.com, surenb@google.com, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, kernel-team@android.com X-Rspamd-Queue-Id: 497DB180009 X-Rspam-User: X-Stat-Signature: qa9ah9abyddr6u11rpkopk9z3acqiitf X-Rspamd-Server: rspam03 X-HE-Tag: 1700587017-580909 X-HE-Meta: U2FsdGVkX18fiSpUGn2IEYLDTHX8YWyPfUh5SnwugHFFJAn+NC3kSEl8usbh8kFEm8k/dqb4ZM2A307inLVjQmM/f+Ud9HbXS72aUD0pxaLhlnipaEMDzl3VL8rI8eAmZo2h/gtesoVAIxclzxHI9Aihli4h3v+fYo+K6VBhzGPGQSUeKbEnCfptX4Y9yg+I91b2ifC9kXsVHyIX2QZtYc3KdaJa9vshPG4adIxtQQBo33vMAtMOLxgCV3YkUpONxl0/U+FekG4dNvm5MUa8CFZNfeX0/0dvBy453y5Q9LFi/2i7J1zSG6JZOIyB6LN89fNOOorO4R7FwPJpkhgVu52KMf888T4+5TlbJgeUy37dVY39tArpJE9eGSmHeUcDx6vXaNvItoNlkehR7PJGbMoIBffDesR3ZXJ8ogh+Xlrz/A/+9oYW1mqhJfAxR4Tm49NS/skGWMch15NXPaccud+Moxv5zk1GpkCfebYZucbl0+qGiGCMBZnFrF/Tr1eTj6cLQHZEdM3DbBF79NA5zVqILAQ0tdH9n94QhE0fhVi93WCCih+08Gze3lbL6wXcXWbRXWP6MUkfgSE9QRIwVcBTqTUfjTqU2pkqhxWm0tVHP4bZEOe0IAPIBl4kHbReAuiJbaBWWewERobskbqn0uxSMJeQNPiv0JqQdkTJbKlhADGSPM9S+GJM1J8+J4xkMpUiKttDvvVrGThBv3RLweAUfeDNNIEdhCz4InotJ8LsV5GeaAz0GpHSUS9J3vpCuXFTfg/LWAH4v0ECDyIf0rDb4Ex5YkUsyn2E8lDhZeLKpWnQnMEmL6/XCJBLlefP8pRPFWyQXvsj1zALL7OxOpLaICfE8EG7kxkP7zIQx2AA/OCYiAr5xeTVWvuaEfEG/SVo6zHiMX26yIwlb0lnVv7GBdNv8JDzmxY0+2Lb4AElRqY5Ot6SSC5VHmEoIeWaw+jFvQ0Wmzkki7SpOx4 4mDUI7+h 9p259wFFdDkMJuYAilpKGgmRV9blsYsnMMLCHoGhEJ2Yfxz68tJW+Ytin/t+Gxmc5MQagh5/HxcyJ4ZR+rI3XN/lb4kDtLU95wickw9YFFRwRMDSvuTMF9/7uXYGOmg5MTwzFMJKBuIWhWDrM7YWBeaBI9IMxoHNF38e2WcoX1VP/BYcqzxCzo/TRpYK0UvpJ6jGx8O/PExuTJ47qh/jJx0YM0mCt+nrjKIqxKhGV1N0MDpLFPW0xK/UNjp6cdNSi6fDydH4jM7re02KdiItBzm54IKwzXWvuUjMrgiyaNnF9gSMAm9uFYuMEknFJDyf4vlTdix5DVLDJk4NXWqkkq2XtvdYsViNVJDqgIL4r+DPSlmb15hVMu+YqyRGRQt6OBEGq1650JHCLUZGnVA5rwW2BD7j2I9o5uVqsFggOm3fpV/zFSnlc0s2gElQ2PW/DebZO9OM02R3y3ZDa68Vwm680PfF3tQF9/IrJXY70rHsdiG/9xRp3+Hzskz5qxgkXLnPCYovwK9Xm01ckpBC07XFsFZKTYhIHMFfcjElnCWLSFj/GWtKKBF8T1CsSEb8MmD6TxGaIFe3xuQVt5KxF6cGkL/wdO4MPKy0ds45zswUUWy0QySgSASt6qY+oW7IHYR7MljImfKGgrfFmW42eucybrQ== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000014, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: Currently each test can specify unique operations using uffd_test_ops, however these operations are per-memory type and not per-test. Add uffd_test_case_ops which each test case can customize for its own needs regardless of the memory type being used. Pre- and post-allocation operations are added, some of which will be used in the next patch to implement test-specific operations like madvise after memory is allocated but before it is accessed. Signed-off-by: Suren Baghdasaryan --- tools/testing/selftests/mm/uffd-common.c | 13 +++++++++++++ tools/testing/selftests/mm/uffd-common.h | 7 +++++++ tools/testing/selftests/mm/uffd-unit-tests.c | 2 ++ 3 files changed, 22 insertions(+) diff --git a/tools/testing/selftests/mm/uffd-common.c b/tools/testing/selftests/mm/uffd-common.c index 583e5a4cc0fd..fb3bbc77fd00 100644 --- a/tools/testing/selftests/mm/uffd-common.c +++ b/tools/testing/selftests/mm/uffd-common.c @@ -17,6 +17,7 @@ bool map_shared; bool test_uffdio_wp = true; unsigned long long *count_verify; uffd_test_ops_t *uffd_test_ops; +uffd_test_case_ops_t *uffd_test_case_ops; static int uffd_mem_fd_create(off_t mem_size, bool hugetlb) { @@ -298,6 +299,12 @@ int uffd_test_ctx_init(uint64_t features, const char **errmsg) unsigned long nr, cpu; int ret; + if (uffd_test_case_ops && uffd_test_case_ops->pre_alloc) { + ret = uffd_test_case_ops->pre_alloc(errmsg); + if (ret) + return ret; + } + ret = uffd_test_ops->allocate_area((void **)&area_src, true); ret |= uffd_test_ops->allocate_area((void **)&area_dst, false); if (ret) { @@ -306,6 +313,12 @@ int uffd_test_ctx_init(uint64_t features, const char **errmsg) return ret; } + if (uffd_test_case_ops && uffd_test_case_ops->post_alloc) { + ret = uffd_test_case_ops->post_alloc(errmsg); + if (ret) + return ret; + } + ret = userfaultfd_open(&features); if (ret) { if (errmsg) diff --git a/tools/testing/selftests/mm/uffd-common.h b/tools/testing/selftests/mm/uffd-common.h index 870776b5a323..774595ee629e 100644 --- a/tools/testing/selftests/mm/uffd-common.h +++ b/tools/testing/selftests/mm/uffd-common.h @@ -90,6 +90,12 @@ struct uffd_test_ops { }; typedef struct uffd_test_ops uffd_test_ops_t; +struct uffd_test_case_ops { + int (*pre_alloc)(const char **errmsg); + int (*post_alloc)(const char **errmsg); +}; +typedef struct uffd_test_case_ops uffd_test_case_ops_t; + extern unsigned long nr_cpus, nr_pages, nr_pages_per_cpu, page_size; extern char *area_src, *area_src_alias, *area_dst, *area_dst_alias, *area_remap; extern int uffd, uffd_flags, finished, *pipefd, test_type; @@ -102,6 +108,7 @@ extern uffd_test_ops_t anon_uffd_test_ops; extern uffd_test_ops_t shmem_uffd_test_ops; extern uffd_test_ops_t hugetlb_uffd_test_ops; extern uffd_test_ops_t *uffd_test_ops; +extern uffd_test_case_ops_t *uffd_test_case_ops; void uffd_stats_report(struct uffd_args *args, int n_cpus); int uffd_test_ctx_init(uint64_t features, const char **errmsg); diff --git a/tools/testing/selftests/mm/uffd-unit-tests.c b/tools/testing/selftests/mm/uffd-unit-tests.c index e7d43c198041..debc423bdbf4 100644 --- a/tools/testing/selftests/mm/uffd-unit-tests.c +++ b/tools/testing/selftests/mm/uffd-unit-tests.c @@ -78,6 +78,7 @@ typedef struct { uffd_test_fn uffd_fn; unsigned int mem_targets; uint64_t uffd_feature_required; + uffd_test_case_ops_t *test_case_ops; } uffd_test_case_t; static void uffd_test_report(void) @@ -185,6 +186,7 @@ uffd_setup_environment(uffd_test_args_t *args, uffd_test_case_t *test, { map_shared = mem_type->shared; uffd_test_ops = mem_type->mem_ops; + uffd_test_case_ops = test->test_case_ops; if (mem_type->mem_flag & (MEM_HUGETLB_PRIVATE | MEM_HUGETLB)) page_size = default_huge_page_size(); From patchwork Tue Nov 21 17:16:38 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Suren Baghdasaryan X-Patchwork-Id: 13463397 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 3E090C61D85 for ; Tue, 21 Nov 2023 17:17:03 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id D23B66B04CD; Tue, 21 Nov 2023 12:17:01 -0500 (EST) Received: by kanga.kvack.org (Postfix, from userid 40) id CB7966B04CE; Tue, 21 Nov 2023 12:17:01 -0500 (EST) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id AB1C86B04CF; Tue, 21 Nov 2023 12:17:01 -0500 (EST) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0014.hostedemail.com [216.40.44.14]) by kanga.kvack.org (Postfix) with ESMTP id 956EF6B04CD for ; Tue, 21 Nov 2023 12:17:01 -0500 (EST) Received: from smtpin09.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay10.hostedemail.com (Postfix) with ESMTP id 6BA1CC0A28 for ; Tue, 21 Nov 2023 17:17:01 +0000 (UTC) X-FDA: 81482616642.09.5B69E65 Received: from mail-yb1-f201.google.com (mail-yb1-f201.google.com [209.85.219.201]) by imf18.hostedemail.com (Postfix) with ESMTP id 7F6B11C001C for ; Tue, 21 Nov 2023 17:16:59 +0000 (UTC) Authentication-Results: imf18.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=nZ7U3VsJ; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf18.hostedemail.com: domain of 3CuZcZQYKCHsrtqdmafnnfkd.bnlkhmtw-lljuZbj.nqf@flex--surenb.bounces.google.com designates 209.85.219.201 as permitted sender) smtp.mailfrom=3CuZcZQYKCHsrtqdmafnnfkd.bnlkhmtw-lljuZbj.nqf@flex--surenb.bounces.google.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1700587019; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=p+YU4xakrSZAEeiCshOnJPmfZHIUA9DyOQyL2tdclCw=; b=o4D3yGF2LiZBI5cDvV2YpLNCUuBS1XW4l+jNA2qZL1p7WVB1T/VelWMZWRUObLb4j1KTFN IzJZV3PPLGhsaclcLRkdCouR7+l/wRtx8F5zDUsMXHtPlz3aWV6k0VC71nnm3RIoeBiX/7 JwvA5IJKg9/lLtOF0O5xdm8wy4AeVls= ARC-Authentication-Results: i=1; imf18.hostedemail.com; dkim=pass header.d=google.com header.s=20230601 header.b=nZ7U3VsJ; dmarc=pass (policy=reject) header.from=google.com; spf=pass (imf18.hostedemail.com: domain of 3CuZcZQYKCHsrtqdmafnnfkd.bnlkhmtw-lljuZbj.nqf@flex--surenb.bounces.google.com designates 209.85.219.201 as permitted sender) smtp.mailfrom=3CuZcZQYKCHsrtqdmafnnfkd.bnlkhmtw-lljuZbj.nqf@flex--surenb.bounces.google.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1700587019; a=rsa-sha256; cv=none; b=A2vR1acqkY5LhPM4yUe88r/3m0rZzPzqFnol0yY09mESqhwXpwAGwUBDXMd2I4KLLvWWx2 Mhz5Z9SmSoTSzMLR/PvJ42Kkd17UMRbI9OwJy8vPhsG+M0HYUkpqzQ449udTqNLQ5JKLb2 fb+J8vA+TT2OSJXMJAWRdpsFpZB0Igg= Received: by mail-yb1-f201.google.com with SMTP id 3f1490d57ef6-da3dd6a72a7so7370947276.0 for ; Tue, 21 Nov 2023 09:16:59 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1700587018; x=1701191818; darn=kvack.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=p+YU4xakrSZAEeiCshOnJPmfZHIUA9DyOQyL2tdclCw=; b=nZ7U3VsJCXJVEod0WsQANYbyCpXud+Kdo3BEbPD8k7ldA1S/O3CqIaxjB0LWHKS/p6 wWi7wPUdhgU/I47MKw8qlxOV3LBcKBLdibiqMdN7Ewo9HLO+dYKHzQspTTowPVLYPqal 98hb2PYzJN4tzMIY2ukRLr6PCi0fFQHQgBGIjh+CxQy9qLfHZ1zE5VvLEcrX74XHqD3F /Ke7nPWTLZSX6hpoNJyiaGKmX+UUIM/xKFZwNPlaQ6C7gyMhtaZmYo0OmnjOi/xzPC9p D5kIVMDDYlJY0pkJsxI+OF+cI9qmJ2w0xQjVq89D5xmla/E7AnS8yY2T1pp2PPKd4qy+ 1jVw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1700587018; x=1701191818; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=p+YU4xakrSZAEeiCshOnJPmfZHIUA9DyOQyL2tdclCw=; b=Cgy7BaPI6SSX2SK/DAS7DgL7nzsujrjv8eTpAvEUKHT8rxgs8laRwtQaoVf978kJ5D ztvCEHF18hh2AtLh3+K9hcq0LbYg6KZQ9/qYfPCCshub+uxMfoxPUqUkVN2xkaHJXNbl wt0Qnnx+fhPEQUrqRRp//vw4Xuxb6NB9GnZXb2IZy40Ja7cMWpyHd8+L3fdnDJsxRlm2 Phj7xNAig6WCIb1ijTGszolTs1WSrhIzLOzISW1WQWBecu2xtr3xGNbKlphDTaneuEjZ 33R1EbhHDL4qtwmzec8UilFW4+LHUvVixjV6aOdw2nS06IBs81sgi75bzl0V+TYBjbQh ufaw== X-Gm-Message-State: AOJu0YwB94YNNnR4s3nnE0MUa/g8jfgTdg+bmKK3GtQHw6rVM0nf8fcD GUmK+KtSPekJUEEXrn5RuOpbCV/yTGs= X-Google-Smtp-Source: AGHT+IGnjdfTVee3DSZlNOy3KhbC3j8nIypXTk1N0NzS6E/A59+6kLC/Rjntyn5b7LkHCGkdnVF0AOMslLI= X-Received: from surenb-desktop.mtv.corp.google.com ([2620:15c:211:201:2045:f6d2:f01d:3fff]) (user=surenb job=sendgmr) by 2002:a25:b84:0:b0:d9a:c27e:5f37 with SMTP id 126-20020a250b84000000b00d9ac27e5f37mr249877ybl.3.1700587018586; Tue, 21 Nov 2023 09:16:58 -0800 (PST) Date: Tue, 21 Nov 2023 09:16:38 -0800 In-Reply-To: <20231121171643.3719880-1-surenb@google.com> Mime-Version: 1.0 References: <20231121171643.3719880-1-surenb@google.com> X-Mailer: git-send-email 2.43.0.rc1.413.gea7ed67945-goog Message-ID: <20231121171643.3719880-6-surenb@google.com> Subject: [PATCH v5 5/5] selftests/mm: add UFFDIO_MOVE ioctl test From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: viro@zeniv.linux.org.uk, brauner@kernel.org, shuah@kernel.org, aarcange@redhat.com, lokeshgidra@google.com, peterx@redhat.com, david@redhat.com, hughd@google.com, mhocko@suse.com, axelrasmussen@google.com, rppt@kernel.org, willy@infradead.org, Liam.Howlett@oracle.com, jannh@google.com, zhangpeng362@huawei.com, bgeffon@google.com, kaleshsingh@google.com, ngeoffray@google.com, jdduke@google.com, surenb@google.com, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, kernel-team@android.com X-Rspam-User: X-Rspamd-Server: rspam12 X-Rspamd-Queue-Id: 7F6B11C001C X-Stat-Signature: yb9ru3p8pfdmimrtrmuuxtjcfxxyb7rt X-HE-Tag: 1700587019-216480 X-HE-Meta: U2FsdGVkX1+JCqZ1Tu3wGEBRx42Z9XJ8R9lfC79RK5PD9NSEdZgmYrg8czJD7l79ChhAzMCKSwFpTDfU8wo+P8IgM4MthftD9MySv2THqZZ35kCRpnBO9GDNbdB7FSsLgtXzU8fINcq9ZNHeUezbjCwq6k/dya3iBVZrwa9kEmDr9eVar+vddG2Qu/JabpfRTgFG/RJQE2pWQDBqY15bhoVj0wID04OI3IFBNjT8OTYfiry/iegxOaTfVhvKlBLyCCr/ruCo/LX29NZP0d9AI7CWLfycdtoinAchneJZHCC5chz81FQUGLh0MGQorPbg4kyMDWCZOLNsdvCvI+KlA2+qV//tA5zO6Ik9STily1vPv36OmyYtWSxtWPXrc2D/N6PydHL0ChgNIynyXExruomy6iaOVzfrZ/hgvAoyvV0bRuGk6t18UhJgUkF2r4lDZ0IAF0pw0K0BsssLcolw16vRfTUR2waHZh8snOzlx0XbFHPS3XaVLvHrEEFHDiRF55nbO0oQT2zhFGfUDJ9stOoQ+V/zdw4AX4jwAF8QdnJhz+u5uc2he224uc2Y78mIb92NqF79Eo+h6zF488tqWCbJY9iha/C2dZwEhllIM9pyexNR3JSqAxBhfLtSVIqGYICpKGKFUcVaaq3R7csC9s6fHt3SCQk4MAUiYIBqs+B7Z5/zANXZtAfdNpDWo236CIiJW4YuN1eYgDCbNlC+4AJRn8DtjoUydFknDTBJjFVZgDWA0L9a9Guoqzjq2Wq3gn0U13ktzVI7KhfYBzNy4vduT53yX5exI8Dac/l1kboUxFNUhrwA1po9MquvNAPZbekpnxUflyZQE6TuVv7qrS0RjI0ZhgVw0TGNftienLIuiTqEwt6d2vt1q64c7akgO5i5nmGvWm/HYX/AEsxYdFjEX0uXLJT9io5g2ekg/3Dw3Sc3/c5x4ZhRV4kbGlSL/5Flctrby/EBM5WaVvr j/Bz7Ok9 wf/EYMpreDK5xZG4oMxGMXxgn5r+CD3CccoKya9/H4IdllvJa55v4J6IQFPTH1TOGj5rw+Q5/U7CvZVz85LufaNRKvLq341g/A1vwK3dupx8iDGASecI5XbDszYRxVnfDTEKo/PqQZZiqxkmwt3fl4pQq5TIlP3oPw+glB85c0ETEknCZLGazynod+n4Eo7JyBt+03WiBxC+jqk3/gHf8Jz113RlqRq77vblVy+9z2v+KB09GFw7SLG7gMVswPyAytUW5hgdAQEAX/rlGesjwBwy9RPw399Za1K35MTdHD5Y04JRQ2OIfX0wjj4Bh2TssItCODvTKeFoQqVD6GsLK7Ve8ywK3Y9IihRm6oELcDWsQ6F64DkpWXWk9PXXS3Jp36s7MzAM+RMinyr2OH8C+7pFqReGft0UvggPLWrF5HGlazEk8M0YQZEda2atMB5sGGtMUip1yE6q2eg2yh53b3hIQr5cmgrn2RVCrzpqRiuF06qBCNHE2hY+bXJambXTw8uJgoEOkQk5wpYPYSFTTF/5xm2fZNTGMfxBwQekMgP57qAEViAfsI3x5pWMhUD0L4OjBBTHVYHA9a7JzRYSgLBt2rYnagb1hbFZXzav5sTZCJtxbZ+suKSu+k3ddaoc0++C/z1OhFBJr61blEz7dGaRBDw== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: Add tests for new UFFDIO_MOVE ioctl which uses uffd to move source into destination buffer while checking the contents of both after the move. After the operation the content of the destination buffer should match the original source buffer's content while the source buffer should be zeroed. Separate tests are designed for PMD aligned and unaligned cases because they utilize different code paths in the kernel. Signed-off-by: Suren Baghdasaryan --- tools/testing/selftests/mm/uffd-common.c | 24 +++ tools/testing/selftests/mm/uffd-common.h | 1 + tools/testing/selftests/mm/uffd-unit-tests.c | 189 +++++++++++++++++++ 3 files changed, 214 insertions(+) diff --git a/tools/testing/selftests/mm/uffd-common.c b/tools/testing/selftests/mm/uffd-common.c index fb3bbc77fd00..b0ac0ec2356d 100644 --- a/tools/testing/selftests/mm/uffd-common.c +++ b/tools/testing/selftests/mm/uffd-common.c @@ -631,6 +631,30 @@ int copy_page(int ufd, unsigned long offset, bool wp) return __copy_page(ufd, offset, false, wp); } +int move_page(int ufd, unsigned long offset, unsigned long len) +{ + struct uffdio_move uffdio_move; + + if (offset + len > nr_pages * page_size) + err("unexpected offset %lu and length %lu\n", offset, len); + uffdio_move.dst = (unsigned long) area_dst + offset; + uffdio_move.src = (unsigned long) area_src + offset; + uffdio_move.len = len; + uffdio_move.mode = UFFDIO_MOVE_MODE_ALLOW_SRC_HOLES; + uffdio_move.move = 0; + if (ioctl(ufd, UFFDIO_MOVE, &uffdio_move)) { + /* real retval in uffdio_move.move */ + if (uffdio_move.move != -EEXIST) + err("UFFDIO_MOVE error: %"PRId64, + (int64_t)uffdio_move.move); + wake_range(ufd, uffdio_move.dst, len); + } else if (uffdio_move.move != len) { + err("UFFDIO_MOVE error: %"PRId64, (int64_t)uffdio_move.move); + } else + return 1; + return 0; +} + int uffd_open_dev(unsigned int flags) { int fd, uffd; diff --git a/tools/testing/selftests/mm/uffd-common.h b/tools/testing/selftests/mm/uffd-common.h index 774595ee629e..cb055282c89c 100644 --- a/tools/testing/selftests/mm/uffd-common.h +++ b/tools/testing/selftests/mm/uffd-common.h @@ -119,6 +119,7 @@ void wp_range(int ufd, __u64 start, __u64 len, bool wp); void uffd_handle_page_fault(struct uffd_msg *msg, struct uffd_args *args); int __copy_page(int ufd, unsigned long offset, bool retry, bool wp); int copy_page(int ufd, unsigned long offset, bool wp); +int move_page(int ufd, unsigned long offset, unsigned long len); void *uffd_poll_thread(void *arg); int uffd_open_dev(unsigned int flags); diff --git a/tools/testing/selftests/mm/uffd-unit-tests.c b/tools/testing/selftests/mm/uffd-unit-tests.c index debc423bdbf4..e4e271511db9 100644 --- a/tools/testing/selftests/mm/uffd-unit-tests.c +++ b/tools/testing/selftests/mm/uffd-unit-tests.c @@ -23,6 +23,9 @@ #define MEM_ALL (MEM_ANON | MEM_SHMEM | MEM_SHMEM_PRIVATE | \ MEM_HUGETLB | MEM_HUGETLB_PRIVATE) +#define ALIGN_UP(x, align_to) \ + ((__typeof__(x))((((unsigned long)(x)) + ((align_to)-1)) & ~((align_to)-1))) + struct mem_type { const char *name; unsigned int mem_flag; @@ -1064,6 +1067,178 @@ static void uffd_poison_test(uffd_test_args_t *targs) uffd_test_pass(); } +static void +uffd_move_handle_fault_common(struct uffd_msg *msg, struct uffd_args *args, + unsigned long len) +{ + unsigned long offset; + + if (msg->event != UFFD_EVENT_PAGEFAULT) + err("unexpected msg event %u", msg->event); + + if (msg->arg.pagefault.flags & + (UFFD_PAGEFAULT_FLAG_WP | UFFD_PAGEFAULT_FLAG_MINOR | UFFD_PAGEFAULT_FLAG_WRITE)) + err("unexpected fault type %llu", msg->arg.pagefault.flags); + + offset = (char *)(unsigned long)msg->arg.pagefault.address - area_dst; + offset &= ~(len-1); + + if (move_page(uffd, offset, len)) + args->missing_faults++; +} + +static void uffd_move_handle_fault(struct uffd_msg *msg, + struct uffd_args *args) +{ + uffd_move_handle_fault_common(msg, args, page_size); +} + +static void uffd_move_pmd_handle_fault(struct uffd_msg *msg, + struct uffd_args *args) +{ + uffd_move_handle_fault_common(msg, args, default_huge_page_size()); +} + +static void +uffd_move_test_common(uffd_test_args_t *targs, unsigned long chunk_size, + void (*handle_fault)(struct uffd_msg *msg, struct uffd_args *args)) +{ + unsigned long nr; + pthread_t uffd_mon; + char c; + unsigned long long count; + struct uffd_args args = { 0 }; + char *orig_area_src, *orig_area_dst; + unsigned long step_size, step_count; + unsigned long src_offs = 0; + unsigned long dst_offs = 0; + + /* Prevent source pages from being mapped more than once */ + if (madvise(area_src, nr_pages * page_size, MADV_DONTFORK)) + err("madvise(MADV_DONTFORK) failure"); + + if (uffd_register(uffd, area_dst, nr_pages * page_size, + true, false, false)) + err("register failure"); + + args.handle_fault = handle_fault; + if (pthread_create(&uffd_mon, NULL, uffd_poll_thread, &args)) + err("uffd_poll_thread create"); + + step_size = chunk_size / page_size; + step_count = nr_pages / step_size; + + if (step_size > page_size) { + char *aligned_src = ALIGN_UP(area_src, chunk_size); + char *aligned_dst = ALIGN_UP(area_dst, chunk_size); + + if (aligned_src != area_src || aligned_dst != area_dst) { + src_offs = (aligned_src - area_src) / page_size; + dst_offs = (aligned_dst - area_dst) / page_size; + step_count--; + } + orig_area_src = area_src; + orig_area_dst = area_dst; + area_src = aligned_src; + area_dst = aligned_dst; + } + + /* + * Read each of the pages back using the UFFD-registered mapping. We + * expect that the first time we touch a page, it will result in a missing + * fault. uffd_poll_thread will resolve the fault by moving source + * page to destination. + */ + for (nr = 0; nr < step_count * step_size; nr += step_size) { + unsigned long i; + + /* Check area_src content */ + for (i = 0; i < step_size; i++) { + count = *area_count(area_src, nr + i); + if (count != count_verify[src_offs + nr + i]) + err("nr %lu source memory invalid %llu %llu\n", + nr + i, count, count_verify[src_offs + nr + i]); + } + + /* Faulting into area_dst should move the page or the huge page */ + for (i = 0; i < step_size; i++) { + count = *area_count(area_dst, nr + i); + if (count != count_verify[dst_offs + nr + i]) + err("nr %lu memory corruption %llu %llu\n", + nr, count, count_verify[dst_offs + nr + i]); + } + + /* Re-check area_src content which should be empty */ + for (i = 0; i < step_size; i++) { + count = *area_count(area_src, nr + i); + if (count != 0) + err("nr %lu move failed %llu %llu\n", + nr, count, count_verify[src_offs + nr + i]); + } + } + if (step_size > page_size) { + area_src = orig_area_src; + area_dst = orig_area_dst; + } + + if (write(pipefd[1], &c, sizeof(c)) != sizeof(c)) + err("pipe write"); + if (pthread_join(uffd_mon, NULL)) + err("join() failed"); + + if (args.missing_faults != step_count || args.minor_faults != 0) + uffd_test_fail("stats check error"); + else + uffd_test_pass(); +} + +static void uffd_move_test(uffd_test_args_t *targs) +{ + uffd_move_test_common(targs, page_size, uffd_move_handle_fault); +} + +static void uffd_move_pmd_test(uffd_test_args_t *targs) +{ + uffd_move_test_common(targs, default_huge_page_size(), + uffd_move_pmd_handle_fault); +} + +static int prevent_hugepages(const char **errmsg) +{ + /* This should be done before source area is populated */ + if (madvise(area_src, nr_pages * page_size, MADV_NOHUGEPAGE)) { + /* Ignore only if CONFIG_TRANSPARENT_HUGEPAGE=n */ + if (errno != EINVAL) { + if (errmsg) + *errmsg = "madvise(MADV_NOHUGEPAGE) failed"; + return -errno; + } + } + return 0; +} + +static int request_hugepages(const char **errmsg) +{ + /* This should be done before source area is populated */ + if (madvise(area_src, nr_pages * page_size, MADV_HUGEPAGE)) { + if (errmsg) { + *errmsg = (errno == EINVAL) ? + "CONFIG_TRANSPARENT_HUGEPAGE is not set" : + "madvise(MADV_HUGEPAGE) failed"; + } + return -errno; + } + return 0; +} + +struct uffd_test_case_ops uffd_move_test_case_ops = { + .post_alloc = prevent_hugepages, +}; + +struct uffd_test_case_ops uffd_move_test_pmd_case_ops = { + .post_alloc = request_hugepages, +}; + /* * Test the returned uffdio_register.ioctls with different register modes. * Note that _UFFDIO_ZEROPAGE is tested separately in the zeropage test. @@ -1141,6 +1316,20 @@ uffd_test_case_t uffd_tests[] = { .mem_targets = MEM_ALL, .uffd_feature_required = 0, }, + { + .name = "move", + .uffd_fn = uffd_move_test, + .mem_targets = MEM_ANON, + .uffd_feature_required = UFFD_FEATURE_MOVE, + .test_case_ops = &uffd_move_test_case_ops, + }, + { + .name = "move-pmd", + .uffd_fn = uffd_move_pmd_test, + .mem_targets = MEM_ANON, + .uffd_feature_required = UFFD_FEATURE_MOVE, + .test_case_ops = &uffd_move_test_pmd_case_ops, + }, { .name = "wp-fork", .uffd_fn = uffd_wp_fork_test,