You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用spaCy提取NER时出现AttributeError该如何解决,如何将NER存入新列

报错原因说明

核心问题是变量类型错误:你错误将spacy输出的文档对象当成了Pandas DataFrame使用,具体问题有两个:

  • 代码执行顺序颠倒,且没有正确读取CSV:你先执行了NDF = nlp(open('tweet_data_clean (3).csv', encoding="utf8").read()),此时还没有加载spacy的nlp模型,且你直接把整个CSV的原始文本传入spacy处理,得到的NDF是spacy.tokens.doc.Doc类型的对象,该类没有Pandas DataFrame才有的apply方法,调用apply自然触发属性不存在的报错。
  • 隐藏的名称错误:你定义的函数名是get_NER_locations,调用时写的是get_NER_location,少了末尾的s,就算解决第一个问题也会触发名称不存在的报错。
修复后的代码示例
import pandas as pd
import spacy

# 先加载spacy模型
nlp = spacy.load('en_core_web_sm')
# 正确读取CSV为Pandas DataFrame
NDF = pd.read_csv('tweet_data_clean (3).csv', encoding="utf8")
loc_labels = ['GPE', 'LOC']
ner_locations = []

def get_NER_locations(row):
    tweet_id = row['id']
    tweet = row['tweet']
    doc = nlp(tweet)
    for ent in doc.ents:
        if ent.label_ in loc_labels:
            ner_locations.append([tweet_id, ent.text, ent.label_, spacy.explain(ent.label_)])

# 遍历DataFrame每行处理,需指定axis=1按行遍历
NDF.apply(lambda row : get_NER_locations(row), axis=1)
ner_NDF = pd.DataFrame(ner_locations, columns=['id', 'ent', 'label', 'label_desc'])
merged_NDF = pd.merge(NDF, ner_NDF, on='id', how='outer')
可选优化方案

如果要实现单条tweet的所有NER结果直接合并到原表的同一行,不需要拆分多行,可以直接在原DataFrame新增列存储列表格式的NER结果:

def get_ner_list(row):
    doc = nlp(row['tweet'])
    return [ent.text for ent in doc.ents if ent.label_ in loc_labels]

NDF['ner_locations'] = NDF.apply(get_ner_list, axis=1)

内容的提问来源于stack exchange,提问作者gloriajones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 19:36:04