如何让Python re.findall()匹配完整单行且不返回多余内容?
问题根因
你之前的代码匹配失败有两个核心原因:
- 正则默认是单行匹配模式,
^和$只会匹配整个字符串的首尾,不会按行识别边界,所以直接把整段输出全部匹配返回了 - 匹配规则没有加入文件名、行首无缩进的过滤条件,没法筛掉clang内置依赖生成的无关记录(也就是输出里带
__debug的缩进行)
另外你调用subprocess时用shell=True传拼接的整串命令存在转义风险,建议改成参数列表的传参方式更稳定。
可直接运行的修正代码
import subprocess as subs import re target_file = "simple.cpp" # 执行clang命令,text=True直接返回字符串格式的输出,无需手动转码 ast_res = subs.run( ["clang", "-Xclang", "-ast-dump", target_file], stdout=subs.PIPE, text=True ) # 正则规则说明: # ^ 匹配行首 # [^|]* 匹配行首到第一个竖线之间无其他竖线,对应行首无前置缩进的特征 # 匹配包含UsingDirectiveDecl关键字、同时包含目标文件名的整行 # re.M 开启多行匹配模式,让^和$识别每一行的首尾边界 match_pattern = re.compile( r"^[^|]*UsingDirectiveDecl.*%s.*$" % re.escape(target_file), flags=re.M ) target_lines = match_pattern.findall(ast_res.stdout) for line in target_lines: print("----------------------------") # 不需要行首尾空白就留strip,要保留原始行尾内容就去掉strip print(line.strip())
如果你使用的Python版本低于3.7,不支持
text=True参数,可以把输出处理部分替换为以下写法:ast_res = subs.run( ["clang", "-Xclang", "-ast-dump", target_file], stdout=subs.PIPE ) ast_content = ast_res.stdout.decode("utf-8") target_lines = match_pattern.findall(ast_content)
运行结果
代码执行后只会输出你需要的目标行,不会带多余内容:
---------------------------- |-UsingDirectiveDecl 0x1e840b8 <simple.cpp:2:1, col:17> col:17 Namespace 0x1378e98 'std'
内容的提问来源于stack exchange,提问作者BillaBong Jr.
相关产品推荐
相关产品推荐

