Sway状态栏Python与Rust版本性能异常及轻量应用测量方法咨询
我为Sway WM(https://swaywm.org)编写了一款状态栏,基本功能是每秒打印系统负载、CPU使用率、内存使用率、时间等信息。
该项目(https://gitlab.com/Yellowhat/statusbar)的初衷是学习CI/CD及高干扰环境下的性能测量方法。
项目最初完全用python编写,后来我又用rust实现了功能完全一致的版本(https://gitlab.com/Yellowhat/statusbar/-/tree/master/tools/rust),两个版本输出的信息完全相同。
通过简化版cachegrind统计每秒执行的指令数I refs结果如下:
python版本:约80万条rust版本:约3.5万条
按该结果我原本认为rust版本远轻于python版本。
由于状态栏会随系统持续运行,我希望明确其对系统的影响,因此编写了如下python脚本,在系统idle状态(1小时不操作电脑,亮度100%)下监控系统状态:
from time import sleep while True: stat = open("/proc/stat").readlines()[0].split()[1:8] power = float(open("/sys/class/power_supply/BAT0/power_now").read().strip()) / 10**6 print(f"{stat},{power}") sleep(1)
/proc/stat返回CPU在不同状态下消耗的时间(jiffies),各字段含义:user:用户态下普通进程的执行时间nice:用户态下调整过nice值的进程执行时间system:内核态进程执行时间idle:空闲时间iowait:等待I/O完成的时间irq:处理硬中断的时间softirq:处理软中断的时间
/sys/class/power_supply/BAT0/power_now返回电池的实时输出功率(单位微瓦)
我对比了4种测试场景:
- tty:开机后直接在tty中运行监控脚本
- sway:开机后启动无状态栏的
sway,打开终端运行监控脚本 - sway barpy:开机后启动搭载
python版状态栏的sway,打开终端运行监控脚本 - sway barrs:开机后启动搭载
rust版状态栏的sway,打开终端运行监控脚本
下图展示了4种场景下SUM(user_t, nice_t, system_t, iowait_t, irq_t, softirq_t) - SUM(user_t0, nice_t0, system_t0, iowait_t0, irq_t0, softirq_t0)(即所有CPU非空闲时间累计值减去测试初始值)随时间的变化趋势:
结果显示tty场景非空闲时间最低,其次是无状态栏的sway场景,但出人意料的是rust版状态栏的非空闲时间远高于python版本。
下图展示了4种场景下功耗随时间的变化趋势:
结果显示tty场景功耗最低,其次是无状态栏的sway场景,rust版功耗略高于后者,python版功耗最高,符合我的预期。
第一张CPU时间图的结果说明该场景下rust版本并不比python版本更轻量。
请问是否有更可靠的方法测量应用的轻量程度?我是否对/proc/stat的内容存在误解?
下图展示了python和rust版本的user_t - user_t0、nice_t - nice_t0、system_t - system_t0随时间的变化趋势:
两者的user态时间相近,但rust版本的system和nice态时间远高于python版本。
按照Peter Cordes的建议我执行了如下命令:
perf stat -a -d -e cpu-cycles,cycles,cycles:u,instructions,instructions:u -r 10 <binary>
结果汇总如下表:
| 版本 | 测试时长 | cpu-cycles | cycles | cycles:u | instructions | instructions:u |
|---|---|---|---|---|---|---|
| python | 0 | 285,248,768 | 285,492,709 | 215,681,759 | 323,852,866 | 278,792,657 |
| python | 60 | 3,038,749,610 | 3,046,770,477 | 1,360,477,472 | 2,425,399,588 | 1,612,134,730 |
| python | 120 | 5,802,965,874 | 5,818,929,489 | 2,496,192,062 | 4,536,305,443 | 2,890,941,664 |
| rust | 0 | 1,165,223 | 1,165,869 | 304,188 | 1,319,168 | 442,565 |
| rust | 60 | 3,791,712,516 | 3,799,515,523 | 1,654,232,537 | 3,355,917,423 | 2,073,096,815 |
| rust | 120 | 7,878,549,570 | 7,897,534,278 | 3,341,698,282 | 6,665,643,064 | 4,144,148,632 |
扣除测试初始值(时长=0时的数值)并除以测试时长后,结果如下:
| 版本 | 测试时长 | cpu-cycles | cycles | cycles:u | instructions | instructions:u |
|---|---|---|---|---|---|---|
| python | 60 | 45,891,681 | 46,021,296 | 19,079,929 | 35,025,779 | 22,222,368 |
| python | 120 | 45,980,976 | 46,111,973 | 19,004,253 | 35,103,771 | 21,767,908 |
| rust | 60 | 63,175,788 | 63,305,828 | 27,565,472 | 55,909,971 | 34,544,238 |
| rust | 120 | 65,644,870 | 65,803,070 | 27,844,951 | 55,536,032 | 34,530,884 |
结果再次显示rust版本每秒消耗的CPU周期和执行的指令数均高于python版本。
$ valgrind \ --tool=cachegrind \ --cachegrind-out-file=/dev/null \ --trace-children=yes \ --I1=32768,8,64 \ --D1=32768,8,64 \ --LL=8388608,16,64 \ --cache-sim=yes \ --branch-sim=yes \ <binary>
运行120秒后的结果如下:
python版本:
I refs: 328,117,776 I1 misses: 5,268,249 LLi misses: 15,274 I1 miss rate: 1.61% LLi miss rate: 0.00% D refs: 135,851,173 (95,116,936 rd + 40,734,237 wr) D1 misses: 4,717,331 ( 4,145,263 rd + 572,068 wr) LLd misses: 167,298 ( 57,266 rd + 110,032 wr) D1 miss rate: 3.5% ( 4.4% + 1.4% ) LLd miss rate: 0.1% ( 0.1% + 0.3% ) LL refs: 9,985,580 ( 9,413,512 rd + 572,068 wr) LL misses: 182,572 ( 72,540 rd + 110,032 wr) LL miss rate: 0.0% ( 0.0% + 0.3% ) Branches: 60,115,083 (56,103,351 cond + 4,011,732 ind) Mispredicts: 5,661,255 ( 4,245,148 cond + 1,416,107 ind) Mispred rate: 9.4% ( 7.6% + 35.3% )
rust版本:
I refs: 100,950,027 I1 misses: 351,835 LLi misses: 5,859 I1 miss rate: 0.35% LLi miss rate: 0.01% D refs: 44,227,512 (24,903,985 rd + 19,323,527 wr) D1 misses: 670,307 ( 341,521 rd + 328,786 wr) LLd misses: 34,962 ( 20,657 rd + 14,305 wr) D1 miss rate: 1.5% ( 1.4% + 1.7% ) LLd miss rate: 0.1% ( 0.1% + 0.1% ) LL refs: 1,022,142 ( 693,356 rd + 328,786 wr) LL misses: 40,821 ( 26,516 rd + 14,305 wr) LL miss rate: 0.0% ( 0.0% + 0.1% ) Branches: 19,868,173 (18,608,630 cond + 1,259,543 ind) Mispredicts: 1,289,209 ( 771,555 cond + 517,654 ind) Mispred rate: 6.5% ( 4.1% + 41.1% )
内容的提问来源于stack exchange,提问作者yellowhat

