Python如何从固定格式文本中提取带pass/fail/warn标记的章节头信息
错误原因
你遇到的报错和功能不生效主要是以下几个问题导致的:
file_content.read().splitlines()执行后没有赋值给变量,读取到的行内容直接被丢弃了,后续循环的file_content还是文件对象而非行列表,且文件读取一次后指针已到末尾,无法再拿到内容list.index()方法找不到目标内容时就会抛出ValueError,你当前的需求不需要计算行号截取区间,目标的「章节名+关键字」内容本身就在匹配到的章节开头行内- 调用逻辑有问题:你定义的是
get_word类,导入后没有实例化就直接调用类方法,变量f没有定义也会报错
正则实现方案
用re.match完全可以实现需求,而且逻辑更稳定,不会受正文内容干扰,完整实现如下:
首先是解析逻辑代码:
import re # 匹配章节开头行的正则规则:>>开头,后跟章节名,最后是可选的三个关键字 SECTION_HEADER_RULE = re.compile(r'^>>\s*(.+?)\s+(pass|fail|warn)\s*$') class SectionParser: def __init__(self, file_path): self.file_path = file_path def get_section_info(self, target_section): # 统一目标章节名的空格格式,避免多空格匹配失败 target = ' '.join(target_section.strip().split()) with open(self.file_path, 'r', encoding='utf-8') as f: for line in f: match_res = SECTION_HEADER_RULE.match(line) if not match_res: continue section_name, keyword = match_res.groups() # 统一格式后匹配章节名 if ' '.join(section_name.split()) == target: return f"{section_name} {keyword}" # 未找到对应章节返回None return None
然后是调用代码:
# 假设上面的代码保存在section_parser.py文件中 from section_parser import SectionParser # 实例化解析器,传入你的文件路径 parser = SectionParser('你的目标文件路径.txt') # 传入要查询的章节名 res = parser.get_section_info('name of section a') if res: print(res)
如果需要同时返回整个章节的完整内容,可以调整成如下逻辑:
def get_full_section(self, target_section): target = ' '.join(target_section.strip().split()) content = [] in_target_section = False with open(self.file_path, 'r', encoding='utf-8') as f: for line in f: if line.startswith('>>'): match_res = SECTION_HEADER_RULE.match(line) if match_res: section_name, _ = match_res.groups() if ' '.join(section_name.split()) == target: in_target_section = True content.append(line) continue # 遇到其他章节开头就结束匹配 if in_target_section: break if in_target_section: content.append(line) # 遇到章节结束标记就停止 if line.strip() == '>>END_SECTION': break return ''.join(content) if content else None
内容的提问来源于stack exchange,提问作者Leiah
相关产品推荐
相关产品推荐

