使用Pyan生成调用图时,如何指定初始函数?
指定起始函数优化Pyan调用图生成
针对大型项目用Pyan生成调用图耗时久的问题,确实可以通过过滤子图或定向分析依赖两种方式实现只生成以指定函数为起点的调用图,以下是具体方案:
方案1:生成完整图后过滤目标子图
这种方式操作简单,先生成完整调用图,再通过广度优先搜索(BFS)提取从目标函数出发的所有可达节点和关联边,适合项目规模不是极端大的场景。
import os import glob import logging from collections import deque from pyan.analyzer import CallGraphVisitor from pyan.graphviz import generate_dot # 配置目标起始函数,格式为「模块名.函数名」 TARGET_START_FUNC = "test_code.main.my_entry_func" # 初始化完整调用图 dirname = os.path.dirname(__file__) filenames = glob(os.path.join(dirname, "test_code/**/*.py"), recursive=True) callgraph = CallGraphVisitor(filenames, logger=logging.getLogger(), root=dirname) # 用BFS找出目标函数可达的所有节点 reachable_nodes = set() queue = deque([TARGET_START_FUNC]) reachable_nodes.add(TARGET_START_FUNC) while queue: current_node = queue.popleft() # 遍历所有边,收集当前节点调用的后续节点 for edge in callgraph.edges: if edge.source == current_node and edge.target not in reachable_nodes: reachable_nodes.add(edge.target) queue.append(edge.target) # 过滤出仅包含可达节点的边 filtered_edges = [e for e in callgraph.edges if e.source in reachable_nodes and e.target in reachable_nodes] # 生成并保存过滤后的调用图(dot格式) dot_content = generate_dot(reachable_nodes, filtered_edges, grouped=False) with open("target_callgraph.dot", "w") as f: f.write(dot_content)
方案2:定向分析依赖(高效版)
如果项目规模极大,生成完整图本身就耗时过长,可以通过递归收集目标函数的依赖文件,只分析必要的代码,从根源减少分析量。这种方式需要处理导入解析逻辑,适合超大型项目:
import os import ast import logging from pyan.analyzer import CallGraphVisitor, FunctionDefVisitor from pyan.graphviz import generate_dot # 配置目标函数的模块和名称 TARGET_MODULE = "test_code.main" TARGET_FUNC_NAME = "my_entry_func" root_dir = os.path.dirname(__file__) # 辅助函数:根据模块名找到对应文件 def get_module_file(module_name, root): return os.path.join(root, module_name.replace(".", "/") + ".py") # 第一步:找到目标函数所在文件并解析其AST target_file = get_module_file(TARGET_MODULE, root_dir) with open(target_file, "r") as f: tree = ast.parse(f.read()) # 提取目标函数的AST节点 func_visitor = FunctionDefVisitor() func_visitor.visit(tree) target_func = next(f for f in func_visitor.functions if f.name == TARGET_FUNC_NAME) # 递归收集所有依赖的函数和文件 visited_funcs = set() visited_files = set() visited_files.add(target_file) class CallVisitor(ast.NodeVisitor): def __init__(self): self.calls = [] def visit_Call(self, node): self.calls.append(node) self.generic_visit(node) def collect_dependencies(func, current_module): func_key = (current_module, func.name) if func_key in visited_funcs: return visited_funcs.add(func_key) # 提取当前函数内的所有调用 call_visitor = CallVisitor() call_visitor.visit(func.node) for call in call_visitor.calls: # 处理同模块内的函数调用 if isinstance(call.func, ast.Name): sub_func_name = call.func.id # 重新解析当前文件找子函数 with open(target_file, "r") as f: current_tree = ast.parse(f.read()) sub_func_visitor = FunctionDefVisitor() sub_func_visitor.visit(current_tree) for sub_func in sub_func_visitor.functions: if sub_func.name == sub_func_name: collect_dependencies(sub_func, current_module) # 处理跨模块的函数调用(简化版,需根据项目导入规则调整) elif isinstance(call.func, ast.Attribute) and isinstance(call.func.value, ast.Name): sub_module_name = f"{TARGET_MODULE.split('.')[0]}.{call.func.value.id}" sub_func_name = call.func.attr sub_module_file = get_module_file(sub_module_name, root_dir) if sub_module_file not in visited_files: visited_files.add(sub_module_file) with open(sub_module_file, "r") as f: sub_tree = ast.parse(f.read()) sub_func_visitor = FunctionDefVisitor() sub_func_visitor.visit(sub_tree) for sub_func in sub_func_visitor.functions: if sub_func.name == sub_func_name: collect_dependencies(sub_func, sub_module_name) # 启动依赖收集 collect_dependencies(target_func, TARGET_MODULE) # 仅用收集到的文件生成调用图 callgraph = CallGraphVisitor(list(visited_files), logger=logging.getLogger(), root=root_dir) # 可选:再次过滤出目标函数为起点的子图(同方案1的BFS逻辑) # ...(此处可复用方案1的过滤代码) # 保存最终调用图 dot_content = generate_dot(callgraph.nodes, callgraph.edges, grouped=False) with open("efficient_target_callgraph.dot", "w") as f: f.write(dot_content)
注意事项
- 方案2的导入解析逻辑是简化版,如果项目使用相对导入、from-import等复杂方式,需要扩展
collect_dependencies函数的处理逻辑。 - 若目标函数调用了内置库函数,可在收集依赖时跳过,避免不必要的分析。
内容的提问来源于stack exchange,提问作者Kyle
相关产品推荐
相关产品推荐

