You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Bert-base-NER构建推特数据集NER模型时遇'list'无'ents'属性错误求助

问题解决:AttributeError: 'list' object has no attribute 'ents'

错误原因是你混淆了Hugging Face NER pipeline和spaCy的API用法:Hugging Face的pipeline("ner")返回的是字典组成的列表,而非带有.ents属性的spaCy Doc对象,直接调用.ents会触发属性错误。

修正后的代码

from transformers import AutoTokenizer, AutoModelForTokenClassification
from transformers import pipeline
import pandas as pd  # 确保导入pandas(若使用DataFrame)

tokenizer = AutoTokenizer.from_pretrained("dslim/bert-base-NER")
model = AutoModelForTokenClassification.from_pretrained("dslim/bert-base-NER")

nlp = pipeline("ner", model=model, tokenizer=tokenizer)

def all_ents(v):
    # 直接遍历pipeline返回的列表,提取实体文本和标签
    return [(ent['word'], ent['entity']) for ent in nlp(v)]

df1 = df.copy()
df1['Entities'] = df['text'].apply(all_ents)

df1.head()

补充说明

  • dslim/bert-base-NER的NER pipeline返回的每个字典包含字段:word(实体文本)、entity(实体标签,如PER、LOC、ORG)、score(预测置信度)、start/end(文本位置索引)。
  • 若需要合并被分词器拆分的子词(比如BERT会把"Apple"拆成"App"和"##le"),可添加子词合并逻辑:
def all_ents(v):
    entities = []
    current_ent = None
    for ent in nlp(v):
        word = ent['word']
        # 处理带"##"前缀的子词
        if word.startswith("##"):
            current_ent = (current_ent[0] + word[2:], current_ent[1])
        else:
            if current_ent is not None:
                entities.append(current_ent)
            current_ent = (word, ent['entity'])
    if current_ent is not None:
        entities.append(current_ent)
    return entities

内容的提问来源于stack exchange,提问作者d_stupido_02

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 09:10:28