From patchwork Mon Mar 3 02:03:23 2025 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Sergey Senozhatsky X-Patchwork-Id: 13998064 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on Received: from ( []) by (Postfix) with ESMTP id C0076C282C6 for ; Mon, 3 Mar 2025 02:25:46 +0000 (UTC) Received: by (Postfix) id 4F0DA280010; Sun, 2 Mar 2025 21:25:46 -0500 (EST) Received: by (Postfix, from userid 40) id 4794B28000F; Sun, 2 Mar 2025 21:25:46 -0500 (EST) X-Delivered-To: Received: by (Postfix, from userid 63042) id 2A54F280010; Sun, 2 Mar 2025 21:25:46 -0500 (EST) X-Delivered-To: Received: from ( []) by (Postfix) with ESMTP id 0710328000F for ; Sun, 2 Mar 2025 21:25:46 -0500 (EST) Received: from (a10.router.float.18 []) by (Postfix) with ESMTP id B478D1409EE for ; Mon, 3 Mar 2025 02:25:45 +0000 (UTC) X-FDA: 83178649050.20.BCAD2D7 Received: from ( []) by (Postfix) with ESMTP id C808940007 for ; Mon, 3 Mar 2025 02:25:43 +0000 (UTC) Authentication-Results:; dkim=pass header.s=google header.b=MSeJf3Eo; dmarc=pass (policy=none); spf=pass ( domain of designates as permitted sender) ARC-Seal: i=1; s=arc-20220608;; t=1740968743; a=rsa-sha256; cv=none; b=pWV48PDXw6wOJHMcSgFWqNWd9wGrqywYn5ww3LCdVTvI/CVH3JAyIvJ7YZSxf2YReTZX7L zMggxD0cI61sbjhjdyfYVor7GXZQgZFHeVvYb2U6uWRvN8tjq9wnVgZoi7OR3hWKtbXt2e HrCVolX9gcCsw8Zl6uveRBLdyghfgL8= ARC-Authentication-Results: i=1;; dkim=pass header.s=google header.b=MSeJf3Eo; dmarc=pass (policy=none); spf=pass ( domain of designates as permitted sender) ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed;; s=arc-20220608; t=1740968743; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=LmmGIMp09Xa9+ajjYrlkAeQFbd0l0inskGKz/hBM/Wo=; b=bf8RzrBjeQAZY3V9qgkEd+RX4NZZzSC2T5MheKaGbPx1now4kyEoSShFBWhFCTFnhAD/Si rSSCM8gCtxk6B1NKUA/y2345tRhj68Yv9h7ua49loVNm4rdwRjJ7wBWSGGYPL/FkX18dhM i6CKY1U0io6lR9PZdP3k4L7n8nYHGLc= Received: by with SMTP id d9443c01a7336-223a3c035c9so9001005ad.1 for ; Sun, 02 Mar 2025 18:25:43 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;; s=google; t=1740968743; x=1741573543;; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=LmmGIMp09Xa9+ajjYrlkAeQFbd0l0inskGKz/hBM/Wo=; b=MSeJf3EosmX3wWwpCDC/RrtsVvYJmDGiZ46AKJhsTZks3AphVXCcmIgtgD+VWBb4Jb TUkAw1mexIK0m3uX/SjSOA1N88lQ77f+C3gaPBNEKdV7BPcRDQnBpBcsNYNaQ/dXW6b/ 19wMQLlzPi8XDtTwmb5ctzqXvh3wCBNKVy+5w= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;; s=20230601; t=1740968743; x=1741573543; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=LmmGIMp09Xa9+ajjYrlkAeQFbd0l0inskGKz/hBM/Wo=; b=bj9lf75Ze2gvAdtDbK9bHxdkdZ2lKUOIYqgA3MHe5F96uWE/uWM6vXeQK3RtuCVi6D Z3C3acHqrCoC8E6jaeFTotouE2MsfKYmEV3WsLwROmQX6qxMeXfYOdoef2WULy0iCc6U e+LXaNH6RTBmIN7f4zZvS7KUcGk49R3jTRltAAAVBvUMWR+Xf1cAqjISA4+VoVPUbdkc qVopTR2G6Q9wYuZSIALF5zvBEequDVd5SdyKfNO4kga+nvG/K2OZsZElyBP2uZEpqBjH +xkjw4tHwIsjwrylfmghE3j1lPNrmrVoL5r6eDqoZW6MYWa5Ux/NSzewAURHYuiJ7NvB aVlQ== X-Forwarded-Encrypted: i=1; AJvYcCWdTZPaJLZZ26frFYevYGjDyykdcUdjaIj/ X-Gm-Message-State: AOJu0YwITKaIgxJdXs4YdkX0XY36mzBQQJbS9jK+A3OEgRBz/N0cgkhc C+Ou/Lc8WZA4n1jGwvYPahCz3UZpCOn6GrRUCdclgTKU/Djyp9+pe1FjgC25yw== X-Gm-Gg: ASbGncsRK5Rth+rj6zxAGZjzT8ei6um8UMRuyl0CcpCmrYfOL+1xvzqOwiHEU8JhAgf LV7cZsWU/iCHGc93hw3K9JMe6sQiCPinmyB6MlLiaGgVyJB8Z2wrjT9/t2YCS1TNXai4ftgS4HT oBNrDBS767ldqJ5GyF2TqCLOnLV3GrQjHi8nZeCxr4+qunPEcPMTNDCJQf/GDMpLKVfWOIzPNkZ dcXekA8rwgKJJyHYPRrAeJARhDWzN9W7ZPHT/OdZvPsZajVzuf2EAUHpiEowVA0ImNpeJVffY1O zAhR7nDY1+miL7q7luvTEi9QOQv5SuTsLsKgGAw2amz5dCo= X-Google-Smtp-Source: AGHT+IE+JByFLuBtDUrc5bcjQDL8kXzJ75Xakggo0iV8YwzjqWoQMMXMd+ywhB7PLbJgopN3ZGVCPQ== X-Received: by 2002:a17:902:e84e:b0:21f:6f33:f96 with SMTP id d9443c01a7336-2234a188d50mr267570275ad.6.1740968742708; Sun, 02 Mar 2025 18:25:42 -0800 (PST) Received: from localhost ([2401:fa00:8f:203:1513:4f61:a4d3:b418]) by with UTF8SMTPSA id d9443c01a7336-22350533a10sm67008045ad.247.2025. (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Sun, 02 Mar 2025 18:25:42 -0800 (PST) From: Sergey Senozhatsky To: Andrew Morton Cc: Yosry Ahmed , Hillf Danton , Kairui Song , Sebastian Andrzej Siewior , Minchan Kim ,,, Sergey Senozhatsky Subject: [PATCH v10 14/19] zsmalloc: introduce new object mapping API Date: Mon, 3 Mar 2025 11:03:23 +0900 Message-ID: <> X-Mailer: git-send-email In-Reply-To: <> References: <> MIME-Version: 1.0 X-Rspam-User: X-Rspamd-Server: rspam03 X-Rspamd-Queue-Id: C808940007 X-Stat-Signature: s34twc5bndcbhwc5q57uqxxd98w5ckjh X-HE-Tag: 1740968743-160016 X-HE-Meta: U2FsdGVkX19eM+6zyYbuv3sdnLmx3hCsb7UL13rBgUU9845sb1XJPY7ay16cExIu1AhrFzmKc1g2MdyOrQTg6wYJmQMEnzLME0rz9nOhxM59GYuQaGZiys3jat3c+Fz6yXyRKK+0U+VkNPh/AEN+m3FG1mCDBWHm1BYnFy6D/y9eIZQNMkIhE0um2GuiwTxa5X8aKkCZJIyKKYlM47M3+Nz6x0F/SUBFNvTeZhAjD94nQn5oMEjoVXzoHnvlKuvpjlRI5If7YwL8hwIJsQdKPYhlRnRiCJcBi5ohHKceCbyBJPVUh8CFuyLu3TKRl+llqEIK0Qes8WASnaOTBPUqrdCLNWRgy1i+nMhnWgklsT/Ps7s14fV8R3y7gjlemTyiNQ/Y/jhOhe1zRWhaS1UuMHqEKdmZGKeNvxqjIfH/Ix/fSZSDoq4KDlji7uXCxjIi46LbBf84RyZMg2qV3e4OXPGlIriopmjprBNCAS3wiNMYQIwwVGPwSLbUBWn0rRNx4puy5pCVqFr+nCpXpLLzhJIY1UWf2FSC5pMCcAymeXJXbNQ3GBCaeDKvjjKnfM+3g7lFJtVE92D7dPm+lQMnh+qJssYj1MYcYoD7OMUIhKOpUF6Ix6EzQtkd/oLkYUBP0DQRzpQbxHdNcG48NjSQBf/7IxXVZ95uhMV/Gd73Jw2yYr8Z7PbPwpfVyC8ssoBhpwKZVKAFNeJR6G+avHgwxGoBrMtMzRzjPZnxOIontMi7u34cSKVk90B0/t1R/V5ilhPm4yB/BolBaPBVatIE91TfytsuHcnh7GEMYSpRGKOyIwbYDpY4ITktVo1Xwbf/YrrK6h38+xEkDBqG9iu02cMMXw6qz8pvYxH5kfOWn1R5g6okoImysoZqbxKvduB2MljLMEZMtOUGASSUBvbRRD+o9OjTJog76S+W1qdFPwDDraxao5rZH2QRbENoHmWjhUeXcUN6a6CmtIKCxz+ L7cNoMDn 5bfc2ftYbH5oE7bvCCZs9yLeAYD3S0OMJ2B4kgwfwkGsuHu7rOIK1PfjaEBzLFyu/gmKRh5Khk/RSZDFlI1agNbYdeIkbatMGcC9MryLs6qyDKt3YIVEQD8Zf1Jftutx0RjiDzGrhSKGJ+M6KOtmNk3cUKJ7f9RN528xAyRwGbH9dq/NmdE9yreYN8IkW911Fa9v4N4MOEfR7i2JjPJRw2L2HE00HOi7KxpKcavWjcsE6jWTH+O+i4W/B7gzALUChNZqBn7FCDQysFXgeFzOTbcVO18N4hGPfbroGf7nrHmWxctW0WWWW5D83Sg== X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: Precedence: bulk X-Loop: List-ID: List-Subscribe: List-Unsubscribe: Current object mapping API is a little cumbersome. First, it's inconsistent, sometimes it returns with page-faults disabled and sometimes with page-faults enabled. Second, and most importantly, it enforces atomicity restrictions on its users. zs_map_object() has to return a liner object address which is not always possible because some objects span multiple physical (non-contiguous) pages. For such objects zsmalloc uses a per-CPU buffer to which object's data is copied before a pointer to that per-CPU buffer is returned back to the caller. This leads to another, final, issue - extra memcpy(). Since the caller gets a pointer to per-CPU buffer it can memcpy() data only to that buffer, and during zs_unmap_object() zsmalloc will memcpy() from that per-CPU buffer to physical pages that object in question spans across. New API splits functions by access mode: - zs_obj_read_begin(handle, local_copy) Returns a pointer to handle memory. For objects that span two physical pages a local_copy buffer is used to store object's data before the address is returned to the caller. Otherwise the object's page is kmap_local mapped directly. - zs_obj_read_end(handle, buf) Unmaps the page if it was kmap_local mapped by zs_obj_read_begin(). - zs_obj_write(handle, buf, len) Copies len-bytes from compression buffer to handle memory (takes care of objects that span two pages). This does not need any additional (e.g. per-CPU) buffers and writes the data directly to zsmalloc pool pages. In terms of performance, on a synthetic and completely reproducible test that allocates fixed number of objects of fixed sizes and iterates over those objects, first mapping in RO then in RW mode: OLD API ======= 3 first results out of 10 369,205,778 instructions # 0.80 insn per cycle 40,467,926 branches # 113.732 M/sec 369,002,122 instructions # 0.62 insn per cycle 40,426,145 branches # 189.361 M/sec 369,036,706 instructions # 0.63 insn per cycle 40,430,860 branches # 204.105 M/sec [..] NEW API ======= 3 first results out of 10 265,799,293 instructions # 0.51 insn per cycle 29,834,567 branches # 170.281 M/sec 265,765,970 instructions # 0.55 insn per cycle 29,829,019 branches # 161.602 M/sec 265,764,702 instructions # 0.51 insn per cycle 29,828,015 branches # 189.677 M/sec [..] T-test on all 10 runs ===================== Difference at 95.0% confidence -1.03219e+08 +/- 55308.7 -27.9705% +/- 0.0149878% (Student's t, pooled s = 58864.4) The old API will stay around until the remaining users switch to the new one. After that we'll also remove zsmalloc per-CPU buffer and CPU hotplug handling. The split of map(RO) and map(WO) into read_{begin/end}/write is suggested by Yosry Ahmed. Suggested-by: Yosry Ahmed Signed-off-by: Sergey Senozhatsky Reviewed-by: Yosry Ahmed --- include/linux/zsmalloc.h | 8 +++ mm/zsmalloc.c | 125 +++++++++++++++++++++++++++++++++++++++ 2 files changed, 133 insertions(+) diff --git a/include/linux/zsmalloc.h b/include/linux/zsmalloc.h index a48cd0ffe57d..7d70983cf398 100644 --- a/include/linux/zsmalloc.h +++ b/include/linux/zsmalloc.h @@ -58,4 +58,12 @@ unsigned long zs_compact(struct zs_pool *pool); unsigned int zs_lookup_class_index(struct zs_pool *pool, unsigned int size); void zs_pool_stats(struct zs_pool *pool, struct zs_pool_stats *stats); + +void *zs_obj_read_begin(struct zs_pool *pool, unsigned long handle, + void *local_copy); +void zs_obj_read_end(struct zs_pool *pool, unsigned long handle, + void *handle_mem); +void zs_obj_write(struct zs_pool *pool, unsigned long handle, + void *handle_mem, size_t mem_len); + #endif diff --git a/mm/zsmalloc.c b/mm/zsmalloc.c index afbd72363731..7566070729ee 100644 --- a/mm/zsmalloc.c +++ b/mm/zsmalloc.c @@ -1362,6 +1362,131 @@ void zs_unmap_object(struct zs_pool *pool, unsigned long handle) } EXPORT_SYMBOL_GPL(zs_unmap_object); +void *zs_obj_read_begin(struct zs_pool *pool, unsigned long handle, + void *local_copy) +{ + struct zspage *zspage; + struct zpdesc *zpdesc; + unsigned long obj, off; + unsigned int obj_idx; + struct size_class *class; + void *addr; + + /* Guarantee we can get zspage from handle safely */ + read_lock(&pool->lock); + obj = handle_to_obj(handle); + obj_to_location(obj, &zpdesc, &obj_idx); + zspage = get_zspage(zpdesc); + + /* Make sure migration doesn't move any pages in this zspage */ + zspage_read_lock(zspage); + read_unlock(&pool->lock); + + class = zspage_class(pool, zspage); + off = offset_in_page(class->size * obj_idx); + + if (off + class->size <= PAGE_SIZE) { + /* this object is contained entirely within a page */ + addr = kmap_local_zpdesc(zpdesc); + addr += off; + } else { + size_t sizes[2]; + + /* this object spans two pages */ + sizes[0] = PAGE_SIZE - off; + sizes[1] = class->size - sizes[0]; + addr = local_copy; + + memcpy_from_page(addr, zpdesc_page(zpdesc), + off, sizes[0]); + zpdesc = get_next_zpdesc(zpdesc); + memcpy_from_page(addr + sizes[0], + zpdesc_page(zpdesc), + 0, sizes[1]); + } + + if (!ZsHugePage(zspage)) + addr += ZS_HANDLE_SIZE; + + return addr; +} +EXPORT_SYMBOL_GPL(zs_obj_read_begin); + +void zs_obj_read_end(struct zs_pool *pool, unsigned long handle, + void *handle_mem) +{ + struct zspage *zspage; + struct zpdesc *zpdesc; + unsigned long obj, off; + unsigned int obj_idx; + struct size_class *class; + + obj = handle_to_obj(handle); + obj_to_location(obj, &zpdesc, &obj_idx); + zspage = get_zspage(zpdesc); + class = zspage_class(pool, zspage); + off = offset_in_page(class->size * obj_idx); + + if (off + class->size <= PAGE_SIZE) { + if (!ZsHugePage(zspage)) + off += ZS_HANDLE_SIZE; + handle_mem -= off; + kunmap_local(handle_mem); + } + + zspage_read_unlock(zspage); +} +EXPORT_SYMBOL_GPL(zs_obj_read_end); + +void zs_obj_write(struct zs_pool *pool, unsigned long handle, + void *handle_mem, size_t mem_len) +{ + struct zspage *zspage; + struct zpdesc *zpdesc; + unsigned long obj, off; + unsigned int obj_idx; + struct size_class *class; + + /* Guarantee we can get zspage from handle safely */ + read_lock(&pool->lock); + obj = handle_to_obj(handle); + obj_to_location(obj, &zpdesc, &obj_idx); + zspage = get_zspage(zpdesc); + + /* Make sure migration doesn't move any pages in this zspage */ + zspage_read_lock(zspage); + read_unlock(&pool->lock); + + class = zspage_class(pool, zspage); + off = offset_in_page(class->size * obj_idx); + + if (off + class->size <= PAGE_SIZE) { + /* this object is contained entirely within a page */ + void *dst = kmap_local_zpdesc(zpdesc); + + if (!ZsHugePage(zspage)) + off += ZS_HANDLE_SIZE; + memcpy(dst + off, handle_mem, mem_len); + kunmap_local(dst); + } else { + /* this object spans two pages */ + size_t sizes[2]; + + off += ZS_HANDLE_SIZE; + sizes[0] = PAGE_SIZE - off; + sizes[1] = mem_len - sizes[0]; + + memcpy_to_page(zpdesc_page(zpdesc), off, + handle_mem, sizes[0]); + zpdesc = get_next_zpdesc(zpdesc); + memcpy_to_page(zpdesc_page(zpdesc), 0, + handle_mem + sizes[0], sizes[1]); + } + + zspage_read_unlock(zspage); +} +EXPORT_SYMBOL_GPL(zs_obj_write); + /** * zs_huge_class_size() - Returns the size (in bytes) of the first huge * zsmalloc &size_class.