You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Stanford Stanza NLP的NER实体叠加到NetworkX词语句法图?

修改NetworkX图适配Stanza NER实体的实现方案

核心思路

先基于Stanza的句法分析构建基础图,再针对两种NER实体(单token LOC、多token PER)分别做节点替换/合并,同时保留正确的句法关系。

具体步骤&代码修改

  1. 提取NER实体映射
    从Stanza的doc.ents里拿到实体的文本、类型以及对应的token/word对象,方便后续匹配图中的节点。

  2. 构建基础句法图
    先按原逻辑添加指定词性(NOUN、PROPN、VERB)的节点,以及对应的句法依存边。

  3. 处理单token实体(Taranto)
    找到对应token的原节点,将其替换为实体名节点,同时把原节点的所有前后边转移到新节点上,保证句法关系不丢失。

  4. 处理多token实体(Ignazio Larussa)
    删除两个原token节点,创建合并后的实体节点,转移原节点的所有关联边;额外确保合并后的节点与动词disse之间存在nsubj依存边(因为原句法分析可能只把其中一个token标记为disse的主语)。

完整代码示例

import networkx as nx
import stanza

# 初始化意大利语Stanza pipeline
nlp = stanza.Pipeline('it', processors='tokenize,mwt,pos,lemma,depparse,ner')
text = "Ignazio Larussa disse che Taranto è una bella città."
doc = nlp(text)

# 初始化有向图
G = nx.DiGraph()

# 第一步:构建基础句法图
for sent in doc.sentences:
    for token in sent.tokens:
        for word in token.words:
            # 只添加指定词性的节点
            if word.upos in ['NOUN', 'PROPN', 'VERB']:
                G.add_node(word.text, pos=word.upos)
                # 添加依存边(从父节点指向当前节点)
                if word.head != 0:
                    head_word = sent.words[word.head - 1]
                    if head_word.upos in ['NOUN', 'PROPN', 'VERB']:
                        G.add_edge(head_word.text, word.text, rel=word.deprel)

# 第二步:提取NER实体信息
ner_entities = []
for ent in doc.ents:
    ner_entities.append({
        'text': ent.text,
        'type': ent.type,
        'words': [word for token in ent.tokens for word in token.words]
    })

# 第三步:处理NER实体,调整图结构
for ent in ner_entities:
    ent_text = ent['text']
    ent_words = ent['words']
    ent_word_texts = [w.text for w in ent_words]

    if len(ent_words) == 1:
        # 处理单token实体(Taranto)
        original_word = ent_words[0]
        if original_word.text in G.nodes:
            # 收集原节点的前后节点和边关系
            preds = list(G.predecessors(original_word.text))
            succs = list(G.successors(original_word.text))
            # 创建实体节点
            G.add_node(ent_text, pos=original_word.upos, ner_type=ent['type'])
            # 转移前驱边
            for pred in preds:
                rel = G[pred][original_word.text]['rel']
                G.add_edge(pred, ent_text, rel=rel)
            # 转移后继边
            for succ in succs:
                rel = G[original_word.text][succ]['rel']
                G.add_edge(ent_text, succ, rel=rel)
            # 删除原节点
            G.remove_node(original_word.text)
    else:
        # 处理多token实体(Ignazio Larussa)
        preds = set()
        succs = set()
        pred_rels = {}
        succ_rels = {}

        # 收集原节点的所有边信息并删除原节点
        for word in ent_words:
            if word.text in G.nodes:
                # 记录前驱节点和边关系
                for p in G.predecessors(word.text):
                    preds.add(p)
                    pred_rels[(p, ent_text)] = G[p][word.text]['rel']
                # 记录后继节点和边关系
                for s in G.successors(word.text):
                    succs.add(s)
                    succ_rels[(ent_text, s)] = G[word.text][s]['rel']
                # 删除原token节点
                G.remove_node(word.text)
        
        # 添加合并后的实体节点
        G.add_node(ent_text, pos='PROPN', ner_type=ent['type'])
        # 转移所有边
        for (p, s), rel in pred_rels.items():
            G.add_edge(p, s, rel=rel)
        for (p, s), rel in succ_rels.items():
            G.add_edge(p, s, rel=rel)
        
        # 确保合并后的实体与disse之间有nsubj边
        if 'disse' in G.nodes and not G.has_edge('disse', ent_text):
            G.add_edge('disse', ent_text, rel='nsubj')

# 验证结果
print("图节点:", G.nodes(data=True))
print("图边:", G.edges(data=True))

关键说明

  • 单token实体直接替换节点,保证原句法关系完整;
  • 多token实体合并时,需要手动补全与核心动词disse的主语边,避免原句法分析只关联单个token导致的边缺失;
  • 所有操作都基于Stanza返回的token/word索引和NER结果,确保实体与图节点的匹配准确。

内容的提问来源于stack exchange,提问作者Robert Alexander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 09:40:44