You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何测量Linux线程上下文切换及线程运行时长?

测量Linux线程上下文切换指标的工具及入门示例

要获取线程加载/卸载CPU的时间戳、运行的CPU核心这些指标,以下是几个实用工具及入门用法:

1. ftrace(内核自带,无需额外安装)

ftrace是Linux内核内置的跟踪工具,能直接捕获调度器的切换事件sched_switch,正好满足需求。

入门步骤:

  1. 跟踪目标进程的线程切换事件:

    # 替换<PID>为你的目标进程ID
    trace-cmd record -e sched_switch -p <PID>
    

    执行后会生成trace.dat文件,停止跟踪按Ctrl+C。

  2. 解析并过滤输出:

    trace-cmd report -i trace.dat | grep -E "<PID>|sched_switch"
    

    输出示例(关键字段说明):

    <...>-1    [001]  85023.400000: sched_switch: prev_comm=sleep prev_tid=1 prev_prio=120 prev_state=R ==> next_comm=sleep next_tid=2 next_prio=120
    <...>-2    [001]  85110.200000: sched_switch: prev_comm=sleep prev_tid=2 prev_prio=120 prev_state=R ==> next_comm=sleep next_tid=1 next_prio=120
    
    • 时间戳:85023.400000(单位:毫秒,可换算为你示例中的小数格式)
    • CPU核心:[001]表示核心1
    • 事件转化:
      • 当next_tid是目标线程ID时,该时间戳就是CPU加载线程的时间
      • 当prev_tid是目标线程ID时,该时间戳就是CPU卸载线程的时间

2. perf(内核自带,性能分析工具)

perf同样可以跟踪调度事件,输出格式更简洁,适合快速分析。

入门步骤:

  1. 记录目标进程的调度事件:

    perf record -e sched:sched_switch -p <PID>
    

    停止按Ctrl+C,生成perf.data。

  2. 查看结构化输出:

    perf script
    

    输出示例:

    sleep  1 [001] 85023.400000: sched:sched_switch: prev_comm=swapper/1 prev_tid=0 prev_prio=120 prev_state=R ==> next_comm=sleep next_tid=1 next_prio=120
    sleep  1 [001] 85110.200000: sched:sched_switch: prev_comm=sleep prev_tid=1 prev_prio=120 prev_state=R ==> next_comm=sleep next_tid=2 next_prio=120
    

    同样通过prev_tid和next_tid区分加载/卸载事件,[001]是CPU核心,时间戳直接提取即可。

3. BCC(基于eBPF,灵活定制输出)

BCC用Python脚本封装eBPF,能直接输出你需要的结构化数据,省去后续格式转换的麻烦,适合需要直接用于可视化的场景。

入门脚本示例:

创建thread_switch.py:

from bcc import BPF

# eBPF程序,跟踪sched_switch事件
bpf_text = """
#include <uapi/linux/sched.h>

BPF_PERF_OUTPUT(events);

struct event {
    u32 pid;
    u32 tid;
    u64 ts;
    char event_type[16];
    u32 cpu;
};

TRACEPOINT_PROBE(sched, sched_switch) {
    struct event e = {};
    struct task_struct *prev = (struct task_struct *)args->prev;
    struct task_struct *next = (struct task_struct *)args->next;

    // 记录卸载事件(prev线程被切换出去)
    e.pid = prev->tgid;
    e.tid = prev->pid;
    e.ts = bpf_ktime_get_ns() / 1000000; // 转换为毫秒
    __builtin_memcpy(&e.event_type, "unload", 6);
    e.cpu = bpf_get_smp_processor_id();
    events.perf_submit(args, &e, sizeof(e));

    // 记录加载事件(next线程被切换进来)
    e.pid = next->tgid;
    e.tid = next->pid;
    __builtin_memcpy(&e.event_type, "load", 4);
    events.perf_submit(args, &e, sizeof(e));

    return 0;
}
"""

b = BPF(text=bpf_text)

# 过滤目标PID(替换为你的进程ID)
TARGET_PID = 1

def print_event(cpu, data, size):
    event = b["events"].event(data)
    if event.pid != TARGET_PID:
        return
    # 转换为你示例中的小数时间格式(秒)
    ts_sec = event.ts / 1000.0
    print(f"{event.pid}\t{event.tid}\t{ts_sec:.4f}\t{'CPU加载该线程' if event.event_type == 'load' else 'CPU卸载该线程'}\t{event.cpu}")

b["events"].open_perf_buffer(print_event)
print("PID\tTID\tTime\t事件\tCPU核心")
while True:
    try:
        b.perf_buffer_poll()
    except KeyboardInterrupt:
        exit()

运行脚本:

sudo python3 thread_switch.py

输出会直接匹配你需要的表格格式,方便后续导入到Excel、Python的pandas中做可视化。

可视化快速示例

用Python的pandas和matplotlib处理BCC输出的数据(假设输出保存为switch_data.csv):

import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv("switch_data.csv", sep="\t")

# 按线程分组,计算每个运行时间段
for tid, group in df.groupby("TID"):
    load_times = group[group["事件"] == "CPU加载该线程"]["Time"]
    unload_times = group[group["事件"] == "CPU卸载该线程"]["Time"]
    cpu_list = group[group["事件"] == "CPU加载该线程"]["CPU核心"].tolist()
    # 绘制每个线程的运行时间段
    for load, unload, cpu in zip(load_times, unload_times, cpu_list):
        plt.barh(tid, unload - load, left=load, label=f"CPU {cpu}", height=0.5)

plt.xlabel("时间(秒)")
plt.ylabel("线程ID(TID)")
plt.legend()
plt.title("线程运行时间段及CPU核心分布")
plt.show()

内容的提问来源于stack exchange,提问作者slip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:31:08