如何将XML标注文本转换为CoNLL格式的IOB2标注数据?
XML标注语料转NER任务IOB2格式CoNLL文件方案
做NER任务预处理时,需要将带实体标注的XML文件转换为CoNLL格式的IOB2标注语料,可通过以下方案实现,仅依赖Python标准库,无需额外安装重型工具。
输入XML结构示例
<doc> Some <tag1>annotated text</tag1> in <tag2>XML</tag2>. </doc>
预期输出CoNLL格式
Some O annotated B-TAG1 text I-TAG1 in O XML B-TAG2 . O
可运行实现代码
import xml.etree.ElementTree as ET import re def xml_to_iob(xml_str): root = ET.fromstring(xml_str) results = [] # 递归遍历节点处理文本和标签 def traverse(node, current_tag=None): # 处理节点内部文本 if node.text: tokens = re.findall(r'\w+|[^\w\s]', node.text.strip()) if tokens: if current_tag is None: for tok in tokens: results.append((tok, 'O')) else: tag_upper = current_tag.upper() results.append((tokens[0], f'B-{tag_upper}')) for tok in tokens[1:]: results.append((tok, f'I-{tag_upper}')) # 处理子节点 for child in node: traverse(child, current_tag=child.tag) # 处理子节点后跟随的文本 if child.tail: tokens = re.findall(r'\w+|[^\w\s]', child.tail.strip()) if tokens: for tok in tokens: results.append((tok, 'O')) traverse(root) # 格式化输出为对齐的CoNLL格式 max_token_len = max(len(tok) for tok, _ in results) + 4 output_lines = [f"{tok.ljust(max_token_len)}{tag}" for tok, tag in results] return '\n'.join(output_lines) # 测试调用 if __name__ == '__main__': test_xml = """<doc> Some <tag1>annotated text</tag1> in <tag2>XML</tag2>. </doc>""" print(xml_to_iob(test_xml))
功能说明
- 内置分词规则可自动拆分单词和标点,符合通用NER语料的分词规范
- 实体标签会自动转为大写,符合CoNLL标注的常用格式
- 兼容嵌套标签的XML标注场景,可直接扩展支持复杂标注需求
内容的提问来源于stack exchange,提问作者coreehi
相关产品推荐
相关产品推荐

