You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python及NLP库从非结构化文本提取特定类别实体?

问题

怎么用Python搭配NLP库从特定语境(比如食谱)的普通文本里,提取出属于特定类别(比如食材)的目标词/实体?举几个例子:

  • 输入Add an onion to a bowl of carrots,要提取出onion和carrots
  • 输入Sprinkle with paprika.,得返回paprika
  • 输入不含食材的句子stir well, and cook an additional minute.,就不返回任何实体

目前用Spacy训练的NER模型泛化能力太差:和训练样本类似的文本能正常识别,但碰到没见过的文本就错得离谱,比如下面的例子:

nlp = spacy.load('trained_model')

# 基本正确的输出
document = nlp('Add flour, mustard, and salt')
[(ent.text, ent.label_) for ent in document.ents]
# >> [('Add flour', 'FOOD'), ('mustard', 'FOOD'), ('salt', 'FOOD')]

# 错误输出:把建筑、汽车、松鼠都识别成食材
document = nlp('I took a building, car and squirrel on the weekend')
[(ent.text, ent.label_) for ent in document.ents]
# >> [('building', 'FOOD'), ('car', 'FOOD'), ('squirrel', 'FOOD')]

# 错误输出:把搅拌、烹饪、额外时间都识别成食材
document = nlp('stir well, and cook an additional minute.')
[(ent.text, ent.label_) for ent in document.ents]
# >> [('stir well', 'FOOD'), ('cook', 'FOOD'), ('additional minute.', 'FOOD')]

不想用那种要把每个名词和 exhaustive 食材列表比对的方法,希望找到能精准提取特定实体,或者在通用文本解析基础上做额外类别分类的可行方案。

可行技术方案

1. 微调领域预训练语言模型

  • 选针对食品/食谱领域预训练好的模型(比如Hugging Face上的bert-base-food,或者用通用大模型比如Llama、GPT-2在食谱数据集上微调)
  • 用公开的标注好的食谱实体数据集(比如RecipeNLG、FoodBase)来微调模型,让它学会从上下文里识别食材实体
  • 示例代码(用Hugging Face Transformers):
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline

# 加载领域预训练模型
tokenizer = AutoTokenizer.from_pretrained("bert-base-food")
model = AutoModelForTokenClassification.from_pretrained("bert-base-food")

# 构建NER处理管道
ner_pipeline = pipeline("ner", model=model, tokenizer=tokenizer, aggregation_strategy="simple")

# 测试效果
result = ner_pipeline("Add an onion to a bowl of carrots")
print([ent['word'] for ent in result if ent['entity_group'] == 'FOOD'])
# >> ['onion', 'carrots']
  • 好处:预训练模型自带强大的上下文理解能力,泛化性比自己小样本训练的Spacy模型好太多

2. 规则+统计模型的混合方案

  • 先用Spacy的通用词性标注(POS)提取所有名词/名词短语,过滤掉非名词类的词
  • 再用轻量级文本分类模型(比如FastText)对提取出的名词做“是不是食材”的二分类
  • 示例步骤:
    1. 用Spacy提取名词短语:
    import spacy
    nlp = spacy.load("en_core_web_sm")
    doc = nlp("Add an onion to a bowl of carrots")
    noun_phrases = [chunk.text for chunk in doc.noun_chunks]
    # >> ['an onion', 'a bowl', 'carrots']
    
    1. 用FastText做分类:
    import fasttext
    # 假设已经训练好一个食材分类模型,训练数据格式类似:__label__FOOD onion\n__label__NON_FOOD bowl...
    model = fasttext.load_model("food_classifier.bin")
    for phrase in noun_phrases:
        pred = model.predict(phrase)
        if pred[0][0] == '__label__FOOD':
            print(phrase)
    # >> 'an onion', 'carrots'
    
  • 好处:既保留规则提取的精准性,又有统计模型的泛化能力,避免纯规则的死板和纯NER的误识别

3. 零样本/少样本实体识别

  • 用支持零样本学习的模型(比如Hugging Face的facebook/bart-large-mnli适配NER任务,或者专门的零样本NER模型)
  • 直接给模型指定目标类别(比如“food ingredient”),不用标注大量训练数据
  • 示例代码(用Hugging Face的零样本NER工具):
from transformers import pipeline

zero_shot_ner = pipeline("zero-shot-ner", model="facebook/bart-large-mnli")

result = zero_shot_ner(
    "Sprinkle with paprika.",
    candidate_labels=["food ingredient", "cooking action", "time duration"]
)
print([ent['word'] for ent in result if ent['entity'] == 'food ingredient'])
# >> ['paprika']
  • 好处:不用标注大量数据,适合快速适配新类别或者样本量少的场景

内容的提问来源于stack exchange,提问作者Riccardo Raffini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 22:47:09