You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何使用BCC解除uprobe挂载比挂载慢得多?求优化方案

问题描述

我开发了一个程序,为5000个函数的入口和出口分别挂载uprobe与uretprobe,挂载过程耗时约30秒,但程序退出时的清理流程耗时长达10分钟以上。

相关C代码:

typedef struct FunctionEvent_t {
  u32 pid;
  u64 timestamp;
  
  u64 func_addr;
  u8 entry;
} FunctionEvent;

BPF_PERF_OUTPUT(traceevents);

static void fill_and_submit(struct pt_regs *ctx, FunctionEvent *event) {
  event->pid = bpf_get_current_pid_tgid();
  event->timestamp = bpf_ktime_get_ns();
  event->func_addr = PT_REGS_IP(ctx);
  traceevents.perf_submit(ctx, event, sizeof(FunctionEvent));
}

int do_entry_point(struct pt_regs *ctx) {
  FunctionEvent event = {.entry = 1};
  fill_and_submit(ctx, &event);
  return 0;
}

int do_exit_point(struct pt_regs *ctx) {
  FunctionEvent event = {.entry = 0};
  fill_and_submit(ctx, &event);
  return 0;
}

相关BCC代码:

from bcc import BPF
bpf_instance = BPF(text=bpf_program)
path = "path/to/exe"

for func_name, func_addr in BPF.get_user_functions_and_addresses(path, ".*"):
    func_name = func_name.decode("utf-8")

    if func_addr in addresses:
        continue

    addresses.add(func_addr)
    try:
        bpf_instance.attach_uprobe(name=self.args.path, sym=func_name, fn_name="do_entry_point")
        bpf_instance.attach_uretprobe(name=self.args.path, sym=func_name, fn_name="do_exit_point")
    except Exception as e:
        print(f"Failed to attach to function {func_name}")

请问:

  1. 是否有方法加速卸载流程?
  2. 挂载与卸载能否并行处理?
  3. 挂载5000个函数是否合理?(我只是在测试系统极限)
回答

1. 加速卸载流程的方法

  • 跳过BCC逐个清理逻辑:BCC Python绑定默认会遍历所有已挂载probe逐个调用detach,10000个probe(5000个函数×2)场景下效率极低。可以直接依赖内核自动回收:退出前关闭所有perf reader(如traceevents对应的读取线程),手动删除bpf_instance引用并触发Python垃圾回收,内核会批量清理所有关联probe,速度远快于逐个卸载。
  • 升级BCC/libbpf版本:新版BCC基于libbpf-bootstrap重构,优化了probe管理与清理逻辑,批量卸载性能有明显提升。
  • 避免资源泄漏:确保所有perf buffer读取线程已终止,无BPF对象引用残留,否则会延迟内核资源回收。

2. 挂载与卸载的并行处理

  • 挂载阶段:完全可以并行。attach_uprobe/attach_uretprobe操作相互独立,只要避免同一函数地址重复挂载,可使用concurrent.futures.ThreadPoolExecutor将函数列表分批次并行挂载,能大幅缩短挂载时间(比如从30秒压缩至数秒)。
  • 卸载阶段:理论上可并行,但需注意线程安全。若自行维护probe句柄(不依赖BCC内部列表),可并行调用detach_uprobe/detach_uretprobe,但需测试内核接口的并发兼容性。不过更高效的方式仍是依赖内核批量回收,而非手动并行卸载。

3. 挂载5000个函数的合理性

从测试系统极限的角度,这完全合理,可验证内核在大规模probe场景下的资源承载能力。但生产环境极不建议:

  • 每个uprobe/uretprobe会占用内核kprobe slot、perf event等资源,10000个probe会增加内核调度与内存开销,可能导致系统性能下降;
  • 挂载/卸载时间成本过高,不适合动态监控场景;
  • 大规模probe易触发内核资源限制(如/proc/sys/kernel/kprobe_max),导致挂载失败。

内容的提问来源于stack exchange,提问作者TreeWater

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 17:31:04