BeautifulSoup处理TEI XML时<body>标签异常包裹文档求助
问题分析与解决
问题根源
你遇到的额外<body>标签包裹整个文档的异常,是因为使用了lxml的HTML解析器处理TEI XML文档。HTML解析器会自动按HTML规则补全结构,将根节点<TEI>强制包裹到<body>中,同时原文档内的<body>标签保留,最终导致重复。
修复方案
将BeautifulSoup的解析器切换为XML模式,严格按照XML规则处理文档,避免HTML解析器的自动补全行为。修改代码中的解析初始化行:
soup = bs(file, 'lxml-xml') # 替换原有的 'lxml'
同时可以简化代码里的冗余变量,优化后的完整函数如下:
def choice_reg(text): with open(f"{text}.xml", "r") as f: file = f.read() # 使用XML解析器处理TEI文档 soup = bs(file, 'lxml-xml') tei_root = soup.find("TEI", recursive=True) # 筛选不含<reg>和<abbr>的<choice>标签 target_choices = [el for el in tei_root.find_all('choice') if not el.find('reg') and not el.find('abbr')] for choice in target_choices: orig = choice.find('orig') reg_string = orig.getText().lower() new_reg = tei_root.new_tag('reg') new_reg.string = reg_string orig.insert_after(new_reg) with open(f"{text.split('.')[0]}2.xml", "w") as f: f.write(str(tei_root))
验证说明
切换为XML解析器后,BeautifulSoup会完整保留TEI文档的原有结构,不会自动注入HTML的<body>标签,同时给<choice>添加小写<reg>的功能可以正常运行。
内容的提问来源于stack exchange,提问作者KWunsch
相关产品推荐
相关产品推荐

