如何用Pandas高效查找CPU从C4唤醒后的首个调度进程
问题描述
我有两个DataFrame:
sched:操作系统调度器事件数据,包含CPU、进程名(Process)、时间戳(Timestamp)字段cstate:CPU C状态切换数据,包含CPU、前序状态(previous)、当前状态(current)、进入时间戳(entry)、退出时间戳(exit)字段
需求是找出导致CPU从C4状态唤醒的首个进程:先过滤cstate中前序状态为'C4'的记录,再针对每个CPU,找到时间戳落在该C4记录的entry和exit区间内的首个进程。目前用嵌套for循环实现,性能极差,想知道有没有无循环的Pandas实现方式。
循环实现的示意代码:
for cpu in range(8): for (i, e) in enumerate(entry_timestamps): sched.loc[sched['CPU'] == cpu].between(e, exit_timestamps[i], inclusive='both').head(1)['process_name']
示例数据
sched表
| CPU | Process | Timestamp |
|---|---|---|
| 0 | foo | 0.0034347 |
| 0 | bar | 0.0036777 |
| 1 | foo | 0.0041122 |
cstate表
| CPU | previous | current | entry | exit |
|---|---|---|---|---|
| 0 | C4 | C0 | 0.0034005 | 0.0043267 |
| 1 | C4 | C0 | 0.0046505 | 0.0053889 |
| 2 | C4 | C0 | 0.0054005 | 0.0055007 |
无循环解决方案
可以用Pandas的merge_asof函数高效实现,它专门用于按时间顺序的近似匹配,性能远高于循环:
步骤1:预处理数据
- 过滤
cstate,只保留从C4唤醒的记录(previous == 'C4'),并按CPU和entry排序 - 对
sched按CPU和Timestamp排序(merge_asof要求右表必须按匹配的时间字段排序)
import pandas as pd # 过滤cstate中从C4唤醒的记录,并排序 cstate_c4 = cstate[cstate['previous'] == 'C4'].sort_values(by=['CPU', 'entry']) # 对sched按CPU和时间戳排序 sched_sorted = sched.sort_values(by=['CPU', 'Timestamp'])
步骤2:用merge_asof匹配首个唤醒进程
merge_asof会按CPU分组,为每个C4唤醒事件匹配第一个时间戳大于等于entry的进程,之后再过滤掉时间戳超过exit的记录:
# 执行近似匹配 matched = pd.merge_asof( cstate_c4, sched_sorted, by='CPU', left_on='entry', right_on='Timestamp', direction='forward' # 找第一个>=entry的时间戳 ) # 过滤掉时间戳超出C4退出时间的记录 result = matched[matched['Timestamp'] <= matched['exit']][['CPU', 'entry', 'exit', 'Process']]
结果说明
对示例数据执行后,result的输出为:
| CPU | entry | exit | Process |
|---|---|---|---|
| 0 | 0.0034005 | 0.0043267 | foo |
- CPU0的C4唤醒事件中,首个符合时间区间的进程是
foo - CPU1的C4唤醒时间(0.0046505)晚于
sched中该CPU的唯一进程时间(0.0041122),无匹配 - CPU2无对应
sched数据,无匹配
性能优势
merge_asof是基于Pandas的向量化实现,时间复杂度为O(n log n)(主要来自排序),相比嵌套循环的O(m*n),在数据量较大时性能提升非常明显。
内容的提问来源于stack exchange,提问作者Felipe Balbi
相关产品推荐
相关产品推荐

