You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用spaCy从文本中提取所有组织名称?现有代码仅提取单个

提取所有组织名称的代码修改方案

你的代码只返回第一个组织名称,是因为函数里的return语句会在找到第一个符合ORG标签的实体后立即结束函数执行。要提取所有组织名称,只需要把找到的实体收集到列表中,最后返回整个列表即可。

修改后的代码

import spacy
nlp = spacy.load("en_core_web_sm")  # 确保已加载spaCy英文模型

def ent(doc):
    org_names = []
    # 遍历文本中的所有命名实体
    for entity in nlp(doc).ents:
        if entity.label_ == "ORG":
            org_names.append(entity.text)
    return org_names

# 测试示例文本
d = "Brock Group (American Industrial Partners) acquires Aegion's Energy Services Businesses"
print(ent(d))

输出结果

运行后会返回所有识别到的组织名称列表:

['Brock Group', 'American Industrial Partners', "Aegion's Energy Services Businesses"]

批量处理数据列(如Pandas列)

如果需要处理Pandas数据框中的文本列,可以用apply方法批量提取:

import pandas as pd

# 假设df是你的数据框,text_column是存储目标文本的列名
df['organizations'] = df['text_column'].apply(ent)

内容的提问来源于stack exchange,提问作者Python-data

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 13:50:21