From patchwork Tue Jan 10 07:42:20 2017 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Patchwork-Submitter: Yi Sun X-Patchwork-Id: 9506629 Return-Path: Received: from mail.wl.linuxfoundation.org (pdx-wl-mail.web.codeaurora.org [172.30.200.125]) by pdx-korg-patchwork.web.codeaurora.org (Postfix) with ESMTP id DC2B5606E1 for ; Tue, 10 Jan 2017 07:45:36 +0000 (UTC) Received: from mail.wl.linuxfoundation.org (localhost [127.0.0.1]) by mail.wl.linuxfoundation.org (Postfix) with ESMTP id CF53A28156 for ; Tue, 10 Jan 2017 07:45:36 +0000 (UTC) Received: by mail.wl.linuxfoundation.org (Postfix, from userid 486) id C442228485; Tue, 10 Jan 2017 07:45:36 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 3.3.1 (2010-03-16) on pdx-wl-mail.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-4.2 required=2.0 tests=BAYES_00, RCVD_IN_DNSWL_MED autolearn=ham version=3.3.1 Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) (using TLSv1.2 with cipher AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by mail.wl.linuxfoundation.org (Postfix) with ESMTPS id 2E09E28156 for ; Tue, 10 Jan 2017 07:45:36 +0000 (UTC) Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.84_2) (envelope-from ) id 1cQr5R-0003lv-Ax; Tue, 10 Jan 2017 07:43:17 +0000 Received: from mail6.bemta3.messagelabs.com ([195.245.230.39]) by lists.xenproject.org with esmtp (Exim 4.84_2) (envelope-from ) id 1cQr5Q-0003lN-6N for xen-devel@lists.xenproject.org; Tue, 10 Jan 2017 07:43:16 +0000 Received: from [85.158.137.68] by server-16.bemta-3.messagelabs.com id 55/B6-03637-39094785; Tue, 10 Jan 2017 07:43:15 +0000 X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrMIsWRWlGSWpSXmKPExsVywNykWHfyhJI Ig3laFt+3TGZyYPQ4/OEKSwBjFGtmXlJ+RQJrRsfx8ywFi2wr/r/5ztTA+NSgi5GTQ0igUmLF psfMILaEAK/EkWUzWCFsf4mWg0eB4lxANQ2MErO+NLOBJNgE1CUef+1hArFFBJQk7q2azARSx Cywn1Fi/vHjYJOEBUIl5vTuYeli5OBgEVCVuLTUCyTMK+AusXfpKiaIBXISJ49NBlvGKeAh0X J1HSPEQe4SzW/vM0LUC0qcnPkEbAwz0N7184RAwswC8hLNW2czT2AUmIWkahZC1SwkVQsYmVc xahSnFpWlFukaWeolFWWmZ5TkJmbm6BoaGOvlphYXJ6an5iQmFesl5+duYgQGZj0DA+MOxqa9 focYJTmYlER5U3RLIoT4kvJTKjMSizPii0pzUosPMcpwcChJ8Fb1A+UEi1LTUyvSMnOAMQKTl uDgURLh3Q6S5i0uSMwtzkyHSJ1iVJQS5w0ESQiAJDJK8+DaYHF5iVFWSpiXkYGBQYinILUoN7 MEVf4VozgHo5IwbzPIFJ7MvBK46a+AFjMBLY60KwZZXJKIkJJqYCzb4ZJzPW2i5aYYgeIzvxq NLr1PmxnOw9vg9t/uvqZKzySZW+JXznf3PfJQMU41rGKe8u18ZFSBvbv7tkXznx54+q9VZmfa 1R/Fye1Ka/fYbdRWX1/y+9S9qzaG78QPrGbayPbxR5JLzrvikrNW6dlGvB/DpyZ0salWLfRj3 v05wcftqjnLEiWW4oxEQy3mouJEAL5kCEjGAgAA X-Env-Sender: yi.y.sun@linux.intel.com X-Msg-Ref: server-11.tower-31.messagelabs.com!1484034191!48884458!2 X-Originating-IP: [192.55.52.115] X-SpamReason: No, hits=0.0 required=7.0 tests= X-StarScan-Received: X-StarScan-Version: 9.1.1; banners=-,-,- X-VirusChecked: Checked Received: (qmail 6047 invoked from network); 10 Jan 2017 07:43:14 -0000 Received: from mga14.intel.com (HELO mga14.intel.com) (192.55.52.115) by server-11.tower-31.messagelabs.com with DHE-RSA-AES256-GCM-SHA384 encrypted SMTP; 10 Jan 2017 07:43:14 -0000 Received: from fmsmga004.fm.intel.com ([10.253.24.48]) by fmsmga103.fm.intel.com with ESMTP; 09 Jan 2017 23:43:14 -0800 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.33,343,1477983600"; d="scan'208";a="211582950" Received: from vmmmba-s2600wft.bj.intel.com ([10.240.193.63]) by fmsmga004.fm.intel.com with ESMTP; 09 Jan 2017 23:43:12 -0800 From: Yi Sun To: xen-devel@lists.xenproject.org Date: Tue, 10 Jan 2017 15:42:20 +0800 Message-Id: <1484034155-4521-2-git-send-email-yi.y.sun@linux.intel.com> X-Mailer: git-send-email 1.9.1 In-Reply-To: <1484034155-4521-1-git-send-email-yi.y.sun@linux.intel.com> References: <1484034155-4521-1-git-send-email-yi.y.sun@linux.intel.com> MIME-Version: 1.0 Cc: wei.liu2@citrix.com, he.chen@linux.intel.com, andrew.cooper3@citrix.com, ian.jackson@eu.citrix.com, Yi Sun , jbeulich@suse.com, chao.p.peng@linux.intel.com Subject: [Xen-devel] [RFC 01/16] docs: create Memory Bandwidth Allocation (MBA) feature document. X-BeenThere: xen-devel@lists.xen.org X-Mailman-Version: 2.1.18 Precedence: list List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xen.org Sender: "Xen-devel" X-Virus-Scanned: ClamAV using ClamSMTP This patch creates MBA feature document in doc/features/. It describes details for MBA. Signed-off-by: Yi Sun --- docs/features/intel_psr_mba.pandoc | 226 +++++++++++++++++++++++++++++++++++++ 1 file changed, 226 insertions(+) create mode 100644 docs/features/intel_psr_mba.pandoc diff --git a/docs/features/intel_psr_mba.pandoc b/docs/features/intel_psr_mba.pandoc new file mode 100644 index 0000000..8c04f01 --- /dev/null +++ b/docs/features/intel_psr_mba.pandoc @@ -0,0 +1,226 @@ +% Intel Memory Bandwidth Allocation (MBA) Feature +% Revision 1.0 + +\clearpage + +# Basics + +---------------- ---------------------------------------------------- + Status: **Tech Preview** + +Architecture(s): Intel x86 + + Component(s): Hypervisor, toolstack + + Hardware: MBA is supported on Skylake Server and beyond +---------------- ---------------------------------------------------- + +# Overview + +The Memory Bandwidth Allocation (MBA) feature provides indirect and approximate +control over memory bandwidth available per-core. This feature provides OS/VMMs +the ability to slow misbehaving apps/VMs or create advanced closed-loop control +system via exposing control over a credit-based throttling mechanism. + +## Terminology + +* CAT Cache Allocation Technology +* COS/CLOS Class of Service +* MSRs Machine Specific Registers +* PSR Intel Platform Shared Resource +* VMM Virtual Machine Monitor +* THRTL Throttle value or delay value + +# User details + +* Feature Enabling: + + Add "psr=mba" to boot line parameter to enable MBA feature. + +* xl interfaces: + + 1. `psr-mba-show [domain-id]`: + + Show system/domain MBA information. + + 2. `psr-mba-set [OPTIONS] domain-id throttling`: + + Set memory bandwidth throttling for domain. + + Options: + '-s': Specify the socket to process, otherwise all sockets are processed. + + Throttling value set in register implies memory bandwidth blocked, i.e. + higher throttling value results in lower bandwidth. The max throttling + value can be got through CPUID. + + The response of the throttling value could be linear mode or non-linear + mode. + + Linear mode: the input precision is defined as 100-(MBA_MAX). For instance, + if the MBA_MAX value is 90, the input precision is 10%. Values not an even + multiple of the precision (e.g., 12%) will be rounded down (e.g., to 10% + delay applied) by HW automatically. + + Non-linear mode: input delay values are powers-of-two from zero to the + MBA_MAX value from CPUID. In this case any values not a power of two will + be rounded down the next nearest power of two by HW automatically. + +# Technical details + +MBA is a member of Intel PSR features, it would share some base PSR +infrastructure in Xen. + +## Hardware perspective + +MBA provides an architectural consistent method to map cores’ to a Class +of Service (COS). This infrastructure will be shared with the previously +introduced CAT technologies. + +Furthermore, MBA also defines a new range MSRs to support specifying a +delay value (Thrtl) per COS, with details below. + ++----------------------------+----------------+ +| MSR (per socket) | Address | ++----------------------------+----------------+ +| IA32_L2_QOS_Ext_BW_Thrtl_0 | 0xD50 | ++----------------------------+----------------+ +| ... | ... | ++----------------------------+----------------+ +| IA32_L2_QOS_Ext_BW_Thrtl_n | 0xD50+n (n<64) | ++----------------------------+----------------+ + +When context switch happens, the COS of VCPU is written to per-thread +MSR `IA32_PQR_ASSOC`, and then hardware enforces bandwidth allocation +according to the throttling value corresponding to the COS. + +## The relationship between MBA and CAT/CDP + +Generally speaking, MBA is completely independent of CAT/CDP, and any +combination may be applied at any time, e.g. enabling MBA with CAT +disabled. + +But it needs to be noticed that MBA shares COS infrastructure with CAT, +although MBA is enumerated by different CPUID leaf from CAT (which +indicates that the max COS of MBA may be different from CAT). + +## Design Overview + +* Core COS/Thrtl association + + When enforcing Memory Bandwidth Allocation, all cores of domains have + the same default COS (COS0) which correspond to the same Thrtl (0). + The default COS is used only in hypervisor and is transparent to tool + stack and user. + + System administrator can change PSR allocation policy at runtime by + tool stack. Since MBA shares COS with CAT/CDP, a COS corresponds to a + 2-tuple, like [CBM, Thrtl] with only-CAT enalbed, when CDP is enable, + the COS corresponds to a 3-tuple, like [Code_CBM, Data_CBM, Thrtl]. If + neither CAT nor CDP is enabled, things would be easier, one COS + corresponds to one Thrtl. + +* VCPU schedule + + This part reuses CAT COS infrastructure. + +* Multi-sockets + + Different sockets may have different MBA ability (like max COS) + although it is consistent on the same socket. So the capability + of per-socket MBA is specified. + +## Implementation Description + +* Hypervisor interfaces: + + 1. Boot line param: "psr=mba" to enable the feature. + + 2. SYSCTL: + - XEN_SYSCTL_PSR_MBA_get_info: Get system MBA information. + + 3. DOMCTL: + - XEN_DOMCTL_PSR_MBA_OP_GET_THRTL: Get Throttling for a domain. + - XEN_DOMCTL_PSR_MBA_OP_SET_THRTL: Set Throttling for a domain. + +* xl interfaces: + + 1. psr-mba-show [domain-id] + Show system/runtime MBA information. + => XEN_SYSCTL_PSR_MBA_get_info/XEN_DOMCTL_PSR_MBA_OP_GET_THRTL + + 2. psr-mba-set [OPTIONS] domain-id throttling + Set bandwidth throttling for a domain. + => XEN_DOMCTL_PSR_MBA_OP_SET_THRTL + +* Key data structure: + + 1. Feature HW info + + ``` + struct psr_mba_hw_info { + unsigned int thrtl_max; + unsigned int cos_max; + unsigned int linear; + }; + + - Member `thrtl_max` + + `thrtl_max` is the max throttling value to be set. + + - Member `cos_max` + + `cos_max` is one of the hardware info of CAT. + + - Member `linear` + + `thrtl_max` means the response of delay value is linear or not. + + As mentioned above, MBA is a member of Intel PSR features, it would + share some base PSR infrastructure in Xen. So, for other data structure + details, please refer 'intel_psr_l2_cat.pandoc'. + +# Limitations + +MBA can only work on HW which enables it (check by CPUID). + +# Testing + +We can execute these commands to verify MBA on different HWs supporting them. + +For example: + root@:~$ xl psr-hwinfo --mba + Memory Bandwidth Allocation (MBA): + Socket ID : 0 + Linear Mode : Enabled + Maximum COS : 7 + Maximum Throttling Value: 90 + Default Throttling Value: 0 + + root@:~$ xl psr-mba-set 1 0xa + + root@:~$ xl psr-mba-show 1 + Socket ID : 0 + Default THRTL : 0 + ID NAME CBM + 1 ubuntu14 0xa + +# Areas for improvement + +N/A + +# Known issues + +N/A + +# References + +"INTEL® RESOURCE DIRECTOR TECHNOLOGY (INTEL® RDT) ALLOCATION FEATURES" [Intel® 64 and IA-32 Architectures Software Developer Manuals, vol3](http://www.intel.com/content/www/us/en/processors/architectures-software-developer-manuals.html) + +# History + +------------------------------------------------------------------------ +Date Revision Version Notes +---------- -------- -------- ------------------------------------------- +2017-01-10 1.0 Xen 4.9 Design document written +---------- -------- -------- -------------------------------------------