如何使用Python提取AWK脚本依赖以生成Graphviz依赖关系图
脚本依赖关系图工具开发说明
我正在使用Python和Graphviz制作可运行于任意机器的全脚本依赖关系图,目前已完成Python脚本与对应依赖的关联、图表生成功能。
当前生成的图表(Python依赖关系图)颜色规则:
- 绿色:非库脚本,无其他脚本依赖它
- 蓝色:库脚本
- 深红色:非库脚本,至少有1个其他脚本依赖它
接下来我计划拓展支持更多语言,首个适配目标为AWK。
核心问题
如何确定AWK脚本中使用的依赖项?
我此前从未接触过AWK,哪怕是实现思路的切入点也十分有用。我已经完成了路径和文件扩展名解析的相关脚本,只需要从AWK脚本中提取出它使用的依赖名称,就可以构建键值对映射。
期望输出格式示例
awk_dict = { 'awk_script_1': ['dependency1', 'dependency2', '...'], 'awk_script_2': ['dependency1', 'dependency2', '...'], ... 'awk_script_n': ['dependency1', 'dependency2', '...'] }
现有Python解析实现参考
应要求附上我已完成的Python脚本解析实现代码:
主逻辑
def main(): """ 生成服务器脚本及其关联关系的图表 """ directory_of_execs = get_executables() generate_graphviz_diagram(parse_script(directory_of_execs))
上述逻辑适用于所有可执行文件,我后续会过滤出参数指定的目标文件,目前运行速度很快,所以暂时没有将过滤逻辑迁移到此处。
可执行文件扫描函数
def get_executables(): """ 生成服务器上所有活跃可执行文件的路径列表 仅抓取当前拥有可执行权限的文件 如果你没有sudo权限,可能会遗漏部分可执行文件,因为你对对应目录或文件的访问权限会被拒绝 返回值:包含所有可访问、非系统文件的活跃可执行文件路径的列表 """ # 允许的可执行文件扩展名 allowed_executables = ('.awk', '.c', '.csh', '.inc', '.ln', '.orig', '.pl', '.pm', '.save', '.sh', '.template', '.py') applicable_dirs = ('/home/', '/Rusr/', '/usr/') exec_info = [] # 此处列出的目录多为系统文件或不重要的生成文件,已确定需要忽略,列表暂不完整 black_listed = ['redhat', 'RHEL', 'local', 'kernel?', 'lib*', 'python*', '__*__', '*system*'] for cleared_dirs in applicable_dirs: for path, dirs, files in os.walk(cleared_dirs, topdown=True, followlinks=False): # 用黑名单过滤当前遍历目录,原地修改生效 dirs[:] = [ d for d in list(dirs) if not any(fnmatch.fnmatch(d, pattern) for pattern in black_listed) if not d.startswith('.') ] # 最终筛选出目标可执行文件 for executable in files: if executable.endswith(allowed_executables): exec_info.append((path, executable)) return exec_info
脚本依赖解析函数
下述函数负责从每个脚本中提取导入模块的信息,并过滤无法编译的脚本:
def parse_script(executables): """ 从输入的脚本中提取所需参数 参数: script_path: 脚本路径字符串 module_str: 仅包含脚本名称和扩展名的字符串 返回值: script_values: 用于生成图表的键值对字典 """ module_container = dict() error_scripts = [] # 因内部错误无法反汇编的脚本 called_scripts = [] # 加入图表的白名单脚本扩展名 ## 注:后续会添加更多支持,目前仅支持Python if ARGS.python: called_scripts.append('.py') for script_path, module_str in executables: # 构建存储脚本信息的字典 script_values = dict() script_values['name'] = module_str[:module_str.rfind('.')].replace('"', '') script_values['extension'] = module_str[module_str.rfind('.'):] script_values['path'] = f'{script_path}/{script_values["name"]}{script_values["extension"]}' # 反汇编并编译脚本 if script_values['extension'] in called_scripts: with open(script_values['path']) as file_pointer: statements = file_pointer.read() try: # 使用dis模块提取单个脚本的指令信息 cat_mod = dis.get_instructions(statements) except Exception as error: # 此处报错并非由当前解析脚本导致,而是被解析的脚本本身存在问题 # 记录异常脚本及对应报错信息 error_scripts.append(f'SCRIPT :: {script_values["name"]}\nPATH ' f':: {script_values["path"]}\n\t' f'ERROR INFO ::\n\t{error}') else: # 仅筛选导入相关的信息 imports = [module for module in cat_mod if 'IMPORT' in module.opname] grouped = defaultdict(list) for imp in imports: grouped[imp.opname].append(imp.argval) script_values['imports'] = grouped # 检查脚本是否已在存储容器中 if script_values['name'] not in module_container: script_values['imports'] = grouped module_container[script_values['name']] = script_values return module_container
我预计需要为每种语言单独编写依赖解析函数,虽然想做一个支持所有语言的通用解析函数,但目前实现难度较大,且不符合代码规范要求。
内容的提问来源于stack exchange,提问作者Bushbaker27
相关产品推荐
相关产品推荐

