使用Bert-base-NER构建推特数据集NER模型时遇'list'无'ents'属性错误求助
问题解决:AttributeError: 'list' object has no attribute 'ents'
错误原因是你混淆了Hugging Face NER pipeline和spaCy的API用法:Hugging Face的pipeline("ner")返回的是字典组成的列表,而非带有.ents属性的spaCy Doc对象,直接调用.ents会触发属性错误。
修正后的代码
from transformers import AutoTokenizer, AutoModelForTokenClassification from transformers import pipeline import pandas as pd # 确保导入pandas(若使用DataFrame) tokenizer = AutoTokenizer.from_pretrained("dslim/bert-base-NER") model = AutoModelForTokenClassification.from_pretrained("dslim/bert-base-NER") nlp = pipeline("ner", model=model, tokenizer=tokenizer) def all_ents(v): # 直接遍历pipeline返回的列表,提取实体文本和标签 return [(ent['word'], ent['entity']) for ent in nlp(v)] df1 = df.copy() df1['Entities'] = df['text'].apply(all_ents) df1.head()
补充说明
- dslim/bert-base-NER的NER pipeline返回的每个字典包含字段:
word(实体文本)、entity(实体标签,如PER、LOC、ORG)、score(预测置信度)、start/end(文本位置索引)。 - 若需要合并被分词器拆分的子词(比如BERT会把"Apple"拆成"App"和"##le"),可添加子词合并逻辑:
def all_ents(v): entities = [] current_ent = None for ent in nlp(v): word = ent['word'] # 处理带"##"前缀的子词 if word.startswith("##"): current_ent = (current_ent[0] + word[2:], current_ent[1]) else: if current_ent is not None: entities.append(current_ent) current_ent = (word, ent['entity']) if current_ent is not None: entities.append(current_ent) return entities
内容的提问来源于stack exchange,提问作者d_stupido_02
相关产品推荐
相关产品推荐

