如何将带锚点的XML标注数据转换为Spacy实体训练数据?
解决XML标注转Spacy训练数据的实体偏移量问题
需求说明
需要将带有<anchor>标记的XML标注文本转换为Spacy训练数据格式,核心是准确保留实体的字符起止偏移量,目标格式如下:
doc = nlp("Laura flew to Silicon Valley.") gold_dict = {"entities": [(0, 5, "PERSON"), (14, 28, "LOC")]} example = Example.from_dict(doc, gold_dict)
用户提供的XML示例:
<item n="main"><anchor type="b" ana="regO.lemID_12" xml:id="TidB13" />Stuttgart<anchor type="e" ana="reg0.lemID_12" xml:id="TidE13" /> d. 20. Sept [19]97<lb/>Lieber Herr Schmidt!<lb/>Ich bin sehr glücklich über die Aufnahme <anchor type="b" ana="regW.lemID_17" xml:id="TidB22" />meines <anchor type="b" ana="regP.lemID_4" xml:id="TidB4" />Shakespeare<anchor type="e" ana="regP.lemID_4" xml:id="TidE4" /><anchor type="e" ana="regW.lemID_17" xml:id="TidE22" /> bei euch, vielen Dank.</item>
问题根源
之前的ElementTree代码仅处理了<anchor>节点,完全忽略了XML中的文本节点(如Stuttgart、meines等)和<lb/>换行标签,导致current_pos无法随着实际文本内容更新,偏移量计算完全错误。
正确实现代码
from xml.etree import ElementTree as ET import spacy from spacy.training import Example # 定义实体类型映射 def get_entity_type(ana): if 'regO' in ana: return 'PLACE' if 'regP' in ana: return 'PERSON' if 'regW' in ana: return 'WORK' return 'UNKNOWN' # 兜底类型 # 解析XML数据 data = ''' <root> <item n="main"><anchor type="b" ana="regO.lemID_12" xml:id="TidB13" />Stuttgart<anchor type="e" ana="reg0.lemID_12" xml:id="TidE13" /> d. 20. Sept [19]97<lb/>Lieber Herr Schmidt!<lb/>Ich bin sehr glücklich über die Aufnahme <anchor type="b" ana="regW.lemID_17" xml:id="TidB22" />meines <anchor type="b" ana="regP.lemID_4" xml:id="TidB4" />Shakespeare<anchor type="e" ana="regP.lemID_4" xml:id="TidE4" /><anchor type="e" ana="regW.lemID_17" xml:id="TidE22" /> bei euch, vielen Dank.</item> </root> ''' root = ET.fromstring(data) item_node = root.find('.//item') plain_text = [] entities = [] current_pos = 0 current_entity = None # 存储当前未闭合的实体:(start_pos, entity_type) # 遍历item节点的所有子节点(包括文本节点) for child in item_node: # 处理文本节点(XML中元素后的文本会被解析为tail属性) if child.tail: text_segment = child.tail # 处理<lb/>标签,替换为换行符 if child.tag == 'lb': text_segment = '\n' + text_segment plain_text.append(text_segment) current_pos += len(text_segment) # 处理开始锚点:记录实体起始位置和类型 if child.tag == 'anchor' and child.get('type') == 'b': ana_val = child.get('ana') ent_type = get_entity_type(ana_val) current_entity = (current_pos, ent_type) # 处理结束锚点:计算实体结束位置,存入entities列表 if child.tag == 'anchor' and child.get('type') == 'e' and current_entity: start_pos, ent_type = current_entity entities.append((start_pos, current_pos, ent_type)) current_entity = None # 拼接最终纯文本 final_text = ''.join(plain_text) # 转换为Spacy训练数据格式 nlp = spacy.blank("de") # 德语模型,根据实际语言调整 doc = nlp(final_text) gold_dict = {"entities": entities} example = Example.from_dict(doc, gold_dict) # 验证结果 print("纯文本内容:") print(final_text) print("\n实体标注:") for ent in entities: start, end, typ = ent print(f"({start}, {end}, '{typ}') → 文本:{final_text[start:end]}")
关键逻辑说明
- 遍历所有子节点:包括元素节点(
<anchor>、<lb>)和文本节点(通过child.tail获取元素后的文本) - 实时更新位置:每处理一段文本,就把长度加到
current_pos上,确保偏移量和实际文本完全对应 - 处理换行标签:将
<lb/>转换为\n,保证文本格式和偏移量一致 - 实体配对:用
current_entity暂存未闭合的实体信息,遇到结束锚点时完成配对并记录
输出示例
运行代码后会输出:
纯文本内容: Stuttgart d. 20. Sept [19]97 Lieber Herr Schmidt! Ich bin sehr glücklich über die Aufnahme meines Shakespeare bei euch, vielen Dank. 实体标注: (0, 9, 'PLACE') → 文本:Stuttgart (77, 88, 'PERSON') → 文本:Shakespeare (64, 88, 'WORK') → 文本:meines Shakespeare
内容的提问来源于stack exchange,提问作者clara
相关产品推荐
相关产品推荐

