You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何构建spaCy抽取管道?公司注册国家抽取遇阻求助

用spaCy抽取企业注册地、资产及股权关联信息的问题

问题背景

待处理的中文文本:

ZZZ LLC是在英国成立的有限责任公司。
XYZ LLC是在英国成立的有限责任公司。XYZ LLC拥有一处位于德国的商业地产,名为‘rentview’。
X先生持有XYZ LLC 21%的股份,剩余79%由ZZZ LLC持有,且ZZZ LLC是XYZ LLC的唯一董事。

期望抽取的结构化JSON结果:

{"name": "XYZ LLC", "type": "ORG", "country": "UK"},
{"name": "ZZZ LLC", "type": "ORG", "country": "UK"},
{"name": "XYZ LLC", "type": "ORG", 
  "owns": [{"name": "rentview", "type": "commercial property", "country": "Germany"}], 
  "owned_by": [
    {"name": "X", "type": "PERSON", "percent": 21},
    {"name":"ZZZ LLC", "type": "ORG", "percent": 79}
  ]
}

当前计划分四步实现:为公司分配注册国家→检测所属资产→检测所有者→生成JSON,但卡在注册国家分配环节,现有实现代码:

Span.set_extension("incorporation_country", default=False)

@Language.component("assign_org_country")
def assign_org_country(doc):
  org_entities = [ent for ent in doc.ents if ent.label_ == "ORG"]
  for ent in org_entities:
    head = ent.root.head
    if head.lemma_ in ['be']:
      for child in head.children:
        if child.dep_ == "attr" and child.text == "company" and child.right_edge.ent_type_ == "GPE":
          ent._.incorporation_country = child
          print(f"country of {ent.text} is {ent._.incorporation_country}")
  return doc

优化思路与实现建议

1. 注册国家抽取的代码优化

当前代码存在明显局限性:

  • 硬匹配child.text == "company",无法兼容中文里“有限公司”“有限责任公司”等变体表述
  • 仅通过child.right_edge定位GPE,逻辑不严谨,中文中“在英国成立”属于介词短语作状语,并非attr的子节点

优化后的注册国家抽取逻辑:

import spacy
from spacy.tokens import Span

nlp = spacy.load("zh_core_web_trf")
Span.set_extension("incorporation_country", default=None)

@Language.component("assign_org_country")
def assign_org_country(doc):
    for ent in doc.ents:
        if ent.label_ == "ORG":
            # 优先匹配"成立/设立"动词的介词宾语
            for token in ent.root.ancestors:
                if token.lemma_ in ["成立", "设立"]:
                    for child in token.children:
                        if child.dep_ == "obl" and child.ent_type_ == "GPE":
                            ent._.incorporation_country = child.text
                            print(f"{ent.text}的注册国家:{ent._.incorporation_country}")
                    break
            # 兼容"是在X成立的公司"的句式
            if not ent._.incorporation_country:
                head = ent.root.head
                if head.lemma_ == "是":
                    for child in head.children:
                        if child.dep_ == "attr":
                            for grandchild in child.subtree:
                                if grandchild.ent_type_ == "GPE":
                                    ent._.incorporation_country = grandchild.text
                                    print(f"{ent.text}的注册国家:{ent._.incorporation_country}")
                                    break
    return doc

# 测试调用
doc = nlp("ZZZ LLC是在英国成立的有限责任公司。XYZ LLC是在英国成立的有限责任公司。")
doc = nlp.add_pipe("assign_org_country", after="ner")(doc)

2. 资产与股权关联的抽取逻辑

资产抽取(公司持有的商业地产)

  • 定位ORG实体作为主语,匹配动词“拥有”/“持有”的宾语
  • 通过文本中的“商业地产”标记资产类型为commercial property
  • 提取附属信息:从“位于德国”中提取GPE实体作为国家,从“名为‘rentview’”中提取资产名称

股权与董事关系抽取

  • 股权:识别“持有X%股份”的句式,提取主语(所有者,PERSON/ORG)、宾语(被持股公司ORG)、百分比(通过数字+%的实体或正则匹配)
  • 董事关系:识别“是XX的董事”的句式,提取主语(董事ORG/PERSON)、宾语(对应公司ORG)

3. 结构化结果整合

  • 用字典存储每个ORG的唯一实例,避免重复条目(比如XYZ LLC只保留一个完整记录)
  • 遍历所有抽取到的关联信息,将注册国家、资产、股权、董事数据合并到对应ORG的字典中
  • 最后将字典转换为指定格式的JSON输出

内容的提问来源于stack exchange,提问作者oktapodia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 01:57:37