如何从Hugging Face的Davlan NER模型提取完整实体名
解决Davlan多语种NER模型实体拆分问题
你遇到的实体拆分问题,可通过在NER pipeline中添加grouped_entities=True参数直接获取完整实体,无需手动合并拆分token。
修改后的代码
from transformers import AutoTokenizer, AutoModelForTokenClassification from transformers import pipeline tokenizer = AutoTokenizer.from_pretrained("Davlan/distilbert-base-multilingual-cased-ner-hrl") model = AutoModelForTokenClassification.from_pretrained("Davlan/distilbert-base-multilingual-cased-ner-hrl") # 启用实体分组,自动合并同类型连续实体 nlp = pipeline("ner", model=model, tokenizer=tokenizer, grouped_entities=True) example = "My name is Johnathan Smith and I work at Apple" ner_results = nlp(example) print(ner_results)
预期输出
[{'entity_group': 'PER', 'score': 0.9997862, 'word': 'Johnathan Smith', 'start': 11, 'end': 26}, {'entity_group': 'ORG', 'score': 0.99870986, 'word': 'Apple', 'start': 41, 'end': 46}]
说明
grouped_entities=True会自动识别IOB格式的B-(实体起始)和I-(实体内部)标签,将同类型的连续拆分token合并为完整实体。- 该参数无需依赖
aggregation_strategy设置,专门针对NER任务的实体合并场景优化,直接返回无拆分的实体名称。
内容的提问来源于stack exchange,提问作者Daniel Wyatt
相关产品推荐
相关产品推荐

