如何构建spaCy抽取管道?公司注册国家抽取遇阻求助
用spaCy抽取企业注册地、资产及股权关联信息的问题
问题背景
待处理的中文文本:
ZZZ LLC是在英国成立的有限责任公司。
XYZ LLC是在英国成立的有限责任公司。XYZ LLC拥有一处位于德国的商业地产,名为‘rentview’。
X先生持有XYZ LLC 21%的股份,剩余79%由ZZZ LLC持有,且ZZZ LLC是XYZ LLC的唯一董事。
期望抽取的结构化JSON结果:
{"name": "XYZ LLC", "type": "ORG", "country": "UK"}, {"name": "ZZZ LLC", "type": "ORG", "country": "UK"}, {"name": "XYZ LLC", "type": "ORG", "owns": [{"name": "rentview", "type": "commercial property", "country": "Germany"}], "owned_by": [ {"name": "X", "type": "PERSON", "percent": 21}, {"name":"ZZZ LLC", "type": "ORG", "percent": 79} ] }
当前计划分四步实现:为公司分配注册国家→检测所属资产→检测所有者→生成JSON,但卡在注册国家分配环节,现有实现代码:
Span.set_extension("incorporation_country", default=False) @Language.component("assign_org_country") def assign_org_country(doc): org_entities = [ent for ent in doc.ents if ent.label_ == "ORG"] for ent in org_entities: head = ent.root.head if head.lemma_ in ['be']: for child in head.children: if child.dep_ == "attr" and child.text == "company" and child.right_edge.ent_type_ == "GPE": ent._.incorporation_country = child print(f"country of {ent.text} is {ent._.incorporation_country}") return doc
优化思路与实现建议
1. 注册国家抽取的代码优化
当前代码存在明显局限性:
- 硬匹配
child.text == "company",无法兼容中文里“有限公司”“有限责任公司”等变体表述 - 仅通过
child.right_edge定位GPE,逻辑不严谨,中文中“在英国成立”属于介词短语作状语,并非attr的子节点
优化后的注册国家抽取逻辑:
import spacy from spacy.tokens import Span nlp = spacy.load("zh_core_web_trf") Span.set_extension("incorporation_country", default=None) @Language.component("assign_org_country") def assign_org_country(doc): for ent in doc.ents: if ent.label_ == "ORG": # 优先匹配"成立/设立"动词的介词宾语 for token in ent.root.ancestors: if token.lemma_ in ["成立", "设立"]: for child in token.children: if child.dep_ == "obl" and child.ent_type_ == "GPE": ent._.incorporation_country = child.text print(f"{ent.text}的注册国家:{ent._.incorporation_country}") break # 兼容"是在X成立的公司"的句式 if not ent._.incorporation_country: head = ent.root.head if head.lemma_ == "是": for child in head.children: if child.dep_ == "attr": for grandchild in child.subtree: if grandchild.ent_type_ == "GPE": ent._.incorporation_country = grandchild.text print(f"{ent.text}的注册国家:{ent._.incorporation_country}") break return doc # 测试调用 doc = nlp("ZZZ LLC是在英国成立的有限责任公司。XYZ LLC是在英国成立的有限责任公司。") doc = nlp.add_pipe("assign_org_country", after="ner")(doc)
2. 资产与股权关联的抽取逻辑
资产抽取(公司持有的商业地产)
- 定位
ORG实体作为主语,匹配动词“拥有”/“持有”的宾语 - 通过文本中的“商业地产”标记资产类型为
commercial property - 提取附属信息:从“位于德国”中提取GPE实体作为国家,从“名为‘rentview’”中提取资产名称
股权与董事关系抽取
- 股权:识别“持有X%股份”的句式,提取主语(所有者,PERSON/ORG)、宾语(被持股公司ORG)、百分比(通过数字+%的实体或正则匹配)
- 董事关系:识别“是XX的董事”的句式,提取主语(董事ORG/PERSON)、宾语(对应公司ORG)
3. 结构化结果整合
- 用字典存储每个
ORG的唯一实例,避免重复条目(比如XYZ LLC只保留一个完整记录) - 遍历所有抽取到的关联信息,将注册国家、资产、股权、董事数据合并到对应
ORG的字典中 - 最后将字典转换为指定格式的JSON输出
内容的提问来源于stack exchange,提问作者oktapodia
相关产品推荐
相关产品推荐

