如何用Python及NLP库从非结构化文本提取特定类别实体?
问题
怎么用Python搭配NLP库从特定语境(比如食谱)的普通文本里,提取出属于特定类别(比如食材)的目标词/实体?举几个例子:
- 输入
Add an onion to a bowl of carrots,要提取出onion和carrots - 输入
Sprinkle with paprika.,得返回paprika - 输入不含食材的句子
stir well, and cook an additional minute.,就不返回任何实体
目前用Spacy训练的NER模型泛化能力太差:和训练样本类似的文本能正常识别,但碰到没见过的文本就错得离谱,比如下面的例子:
nlp = spacy.load('trained_model') # 基本正确的输出 document = nlp('Add flour, mustard, and salt') [(ent.text, ent.label_) for ent in document.ents] # >> [('Add flour', 'FOOD'), ('mustard', 'FOOD'), ('salt', 'FOOD')] # 错误输出:把建筑、汽车、松鼠都识别成食材 document = nlp('I took a building, car and squirrel on the weekend') [(ent.text, ent.label_) for ent in document.ents] # >> [('building', 'FOOD'), ('car', 'FOOD'), ('squirrel', 'FOOD')] # 错误输出:把搅拌、烹饪、额外时间都识别成食材 document = nlp('stir well, and cook an additional minute.') [(ent.text, ent.label_) for ent in document.ents] # >> [('stir well', 'FOOD'), ('cook', 'FOOD'), ('additional minute.', 'FOOD')]
不想用那种要把每个名词和 exhaustive 食材列表比对的方法,希望找到能精准提取特定实体,或者在通用文本解析基础上做额外类别分类的可行方案。
可行技术方案
1. 微调领域预训练语言模型
- 选针对食品/食谱领域预训练好的模型(比如Hugging Face上的
bert-base-food,或者用通用大模型比如Llama、GPT-2在食谱数据集上微调) - 用公开的标注好的食谱实体数据集(比如RecipeNLG、FoodBase)来微调模型,让它学会从上下文里识别食材实体
- 示例代码(用Hugging Face Transformers):
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline # 加载领域预训练模型 tokenizer = AutoTokenizer.from_pretrained("bert-base-food") model = AutoModelForTokenClassification.from_pretrained("bert-base-food") # 构建NER处理管道 ner_pipeline = pipeline("ner", model=model, tokenizer=tokenizer, aggregation_strategy="simple") # 测试效果 result = ner_pipeline("Add an onion to a bowl of carrots") print([ent['word'] for ent in result if ent['entity_group'] == 'FOOD']) # >> ['onion', 'carrots']
- 好处:预训练模型自带强大的上下文理解能力,泛化性比自己小样本训练的Spacy模型好太多
2. 规则+统计模型的混合方案
- 先用Spacy的通用词性标注(POS)提取所有名词/名词短语,过滤掉非名词类的词
- 再用轻量级文本分类模型(比如FastText)对提取出的名词做“是不是食材”的二分类
- 示例步骤:
- 用Spacy提取名词短语:
import spacy nlp = spacy.load("en_core_web_sm") doc = nlp("Add an onion to a bowl of carrots") noun_phrases = [chunk.text for chunk in doc.noun_chunks] # >> ['an onion', 'a bowl', 'carrots']- 用FastText做分类:
import fasttext # 假设已经训练好一个食材分类模型,训练数据格式类似:__label__FOOD onion\n__label__NON_FOOD bowl... model = fasttext.load_model("food_classifier.bin") for phrase in noun_phrases: pred = model.predict(phrase) if pred[0][0] == '__label__FOOD': print(phrase) # >> 'an onion', 'carrots' - 好处:既保留规则提取的精准性,又有统计模型的泛化能力,避免纯规则的死板和纯NER的误识别
3. 零样本/少样本实体识别
- 用支持零样本学习的模型(比如Hugging Face的
facebook/bart-large-mnli适配NER任务,或者专门的零样本NER模型) - 直接给模型指定目标类别(比如“food ingredient”),不用标注大量训练数据
- 示例代码(用Hugging Face的零样本NER工具):
from transformers import pipeline zero_shot_ner = pipeline("zero-shot-ner", model="facebook/bart-large-mnli") result = zero_shot_ner( "Sprinkle with paprika.", candidate_labels=["food ingredient", "cooking action", "time duration"] ) print([ent['word'] for ent in result if ent['entity'] == 'food ingredient']) # >> ['paprika']
- 好处:不用标注大量数据,适合快速适配新类别或者样本量少的场景
内容的提问来源于stack exchange,提问作者Riccardo Raffini
相关产品推荐
相关产品推荐

