Intel-KVM L1D flush异常:未触发及触发后卡顿求助
Intel KVM L1D Flush功能测试问题排查求助
测试背景
正在测试Intel KVM的L1D flush功能性能,对应处理逻辑位于内核源码arch/x86/kvm/vmx/vmx.c的vmx_l1d_flush函数,已添加printk跟踪执行流程。
修改后的vmx_l1d_flush代码
/* * Software based L1D cache flush which is used when microcode providing * the cache control MSR is not loaded. * * The L1D cache is 32 KiB on Nehalem and later microarchitectures, but to * flush it is required to read in 64 KiB because the replacement algorithm * is not exactly LRU. This could be sized at runtime via topology * information but as all relevant affected CPUs have 32KiB L1D cache size * there is no point in doing so. */ static noinstr void vmx_l1d_flush(struct kvm_vcpu *vcpu) { int size = PAGE_SIZE << L1D_CACHE_ORDER; printk(KERN_INFO "vmx_l1d_flush triggered!\n"); /* * This code is only executed when the flush mode is 'cond' or * 'always' */ if (static_branch_likely(&vmx_l1d_flush_cond)) { bool flush_l1d; /* * Clear the per-vcpu flush bit, it gets set again * either from vcpu_run() or from one of the unsafe * VMEXIT handlers. */ flush_l1d = vcpu->arch.l1tf_flush_l1d; // Make every flush happen vcpu->arch.l1tf_flush_l1d = true; /* * Clear the per-cpu flush bit, it gets set again from * the interrupt handlers. */ flush_l1d |= kvm_get_cpu_l1tf_flush_l1d(); kvm_clear_cpu_l1tf_flush_l1d(); if (!flush_l1d) return; } vcpu->stat.l1d_flush++; if (static_cpu_has(X86_FEATURE_FLUSH_L1D)) { native_wrmsrl(MSR_IA32_FLUSH_CMD, L1D_FLUSH); return; } asm volatile( /* First ensure the pages are in the TLB */ "xorl %%eax, %%eax\n" ".Lpopulate_tlb:\n\t" "movzbl (%[flush_pages], %%" _ASM_AX "), %%ecx\n\t" "addl $4096, %%eax\n\t" "cmpl %%eax, %[size]\n\t" "jne .Lpopulate_tlb\n\t" "xorl %%eax, %%eax\n\t" "cpuid\n\t" /* Now fill the cache */ "xorl %%eax, %%eax\n" ".Lfill_cache:\n" "movzbl (%[flush_pages], %%" _ASM_AX "), %%ecx\n\t" "addl $64, %%eax\n\t" "cmpl %%eax, %[size]\n\t" "jne .Lfill_cache\n\t" "lfence\n" :: [flush_pages] "r" (vmx_l1d_flush_pages), [size] "r" (size) : "eax", "ebx", "ecx", "edx"); printk(KERN_INFO "vmx_l1d_flush ended!\n"); }
函数调用点
该函数在vmx_vcpu_enter_exit中被调用,相关代码片段如下:
static noinstr void vmx_vcpu_enter_exit(struct kvm_vcpu *vcpu, struct vcpu_vmx *vmx, unsigned long flags) { guest_state_enter_irqoff(); /* L1D Flush includes CPU buffer clear to mitigate MDS */ if (static_branch_unlikely(&vmx_l1d_should_flush)) vmx_l1d_flush(vcpu); else if (static_branch_unlikely(&mds_user_clear)) mds_clear_cpu_buffers(); else if (static_branch_unlikely(&mmio_stale_data_clear) && kvm_arch_has_assigned_device(vcpu->kvm)) mds_clear_cpu_buffers(); vmx_disable_fb_clear(vmx); if (vcpu->arch.cr2 != native_read_cr2()) native_write_cr2(vcpu->arch.cr2); vmx->fail = __vmx_vcpu_run(vmx, (unsigned long *)&vcpu->arch.regs, flags); vcpu->arch.cr2 = native_read_cr2(); .... (Skipping the rest of the lines since they're not the issue) }
遇到的问题
- 模块参数不生效:根据内核文档设置
kvm_intel模块参数vmentry_l1d_flush=always后,vmx_l1d_flush函数未触发,查看/sys/modules/kvm_intel/parameters/vmentry_l1d_flush仍显示"not required"。 - 手动触发后卡死:修改代码绕过检查强制触发
vmx_l1d_flush后,函数卡在汇编代码的/* Now fill the cache */注释之前,虚拟机黑屏无法进入BIOS。测试使用CPU为Intel(R) Xeon(R) Platinum 8358P @ 2.60GHz。
排查思路建议
- CPU特性与微码检查:Xeon Platinum 8358P属于Ice Lake架构,原生支持硬件L1D flush(
X86_FEATURE_FLUSH_L1D),先确认/proc/cpuinfo中是否存在flush_l1d标志,同时检查系统微码版本是否加载了支持MSR_IA32_FLUSH_CMD的版本——若硬件flush可用,内核会直接走硬件路径而非软件flush,这可能是你看不到软件路径执行的原因。 - 模块参数生效条件排查:
vmentry_l1d_flush的实际状态由内核根据CPU漏洞(如L1TF、MDS)和硬件支持自动判定,查看内核启动日志(dmesg)中KVM初始化阶段的L1TF、MDS相关输出,确认vmx_l1d_should_flush静态分支的触发条件是否满足。 noinstr函数限制检查:vmx_l1d_flush标记为noinstr,该上下文禁止调用可能触发调度、页错误或内存分配的函数,你添加的printk可能违反了这个限制,导致卡死。建议替换为trace_printk(更适合无仪器上下文)或移除printk后再测试。- 软件flush内存合法性检查:确认
vmx_l1d_flush_pages指针是否指向合法、对齐的物理内存区域——软件flush需要遍历这段内存,若内存非法会触发页错误,在noinstr上下文下会直接导致系统卡死或panic。 - 循环逻辑验证:Ice Lake架构的L1D缓存大小为48KiB,而原代码假设为32KiB,检查
L1D_CACHE_ORDER的定义是否适配当前CPU,计算出的size是否正确,避免循环越界导致异常。 - 低级别调试:使用
asm volatile("int3")在关键位置插入断点,配合kgdb调试;或利用内核ftrace跟踪函数执行流程,定位卡死的具体指令。
内容的提问来源于stack exchange,提问作者De Santa Michell
相关产品推荐
相关产品推荐

