You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求协助为Spacey模型添加GPA命名实体(简历数据集无该实体)

为SpaCy模型添加GPA命名实体

方法1:规则匹配快速实现

GPA有固定格式(如3.8/4.0、95%、4.0),用SpaCy的Matcher可以快速匹配并标记实体:

  • 加载基础模型并初始化Matcher
  • 定义覆盖常见GPA格式的匹配规则
  • 为匹配到的文本添加GPA实体标签

示例代码:

import spacy
from spacy.matcher import Matcher

# 加载预训练模型
nlp = spacy.load("en_core_web_sm")
matcher = Matcher(nlp.vocab)

# 定义GPA匹配模式,覆盖主流格式
gpa_patterns = [
    # X.X/X.X 格式(如3.7/4.0)
    [{"IS_DIGIT": True}, {"ORTH": "."}, {"IS_DIGIT": True}, {"ORTH": "/"}, {"IS_DIGIT": True}, {"ORTH": "."}, {"IS_DIGIT": True}],
    # X.X 格式(如4.0)
    [{"IS_DIGIT": True}, {"ORTH": "."}, {"IS_DIGIT": True}],
    # XX/XX 百分制格式(如92/100)
    [{"IS_DIGIT": True}, {"ORTH": "/"}, {"IS_DIGIT": True}],
    # XX% 格式(如95%)
    [{"IS_DIGIT": True}, {"ORTH": "%"}]
]

# 注册匹配规则,命名为GPA
matcher.add("GPA", gpa_patterns)

# 测试处理简历文本
doc = nlp("My college GPA is 3.8/4.0, which translates to 95%.")
matches = matcher(doc)

# 将匹配结果转为实体
for match_id, start, end in matches:
    span = doc[start:end]
    # 合并原有实体与新标记的GPA实体
    doc.ents = list(doc.ents) + [spacy.tokens.Span(doc, start, end, label="GPA")]

# 输出验证
for ent in doc.ents:
    print(f"{ent.text} → {ent.label_}")

方法2:训练自定义NER模型(复杂场景适用)

如果简历中GPA的表达形式多样,规则匹配覆盖不全,可以通过标注数据训练自定义实体模型:

  1. 准备标注好的训练数据,示例格式如下:
TRAIN_DATA = [
    ("I graduated with a GPA of 3.9/4.0", {"entities": [(27, 34, "GPA")]}),
    ("Undergraduate GPA: 92/100", {"entities": [(21, 26, "GPA")]}),
    ("Perfect 4.0 GPA on record", {"entities": [(8, 11, "GPA")]})
]
  1. 加载基础模型,添加GPA标签并训练:
import spacy
from spacy.training import Example

nlp = spacy.load("en_core_web_sm")
ner = nlp.get_pipe("ner")
# 新增GPA实体标签
ner.add_label("GPA")

# 仅训练NER管道,禁用其他管道提升效率
other_pipes = [pipe for pipe in nlp.pipe_names if pipe != "ner"]
with nlp.disable_pipes(*other_pipes):
    optimizer = nlp.begin_training()
    # 迭代训练10轮
    for itn in range(10):
        losses = {}
        for text, annotations in TRAIN_DATA:
            doc = nlp.make_doc(text)
            example = Example.from_dict(doc, annotations)
            nlp.update([example], sgd=optimizer, losses=losses)
        print(f"迭代 {itn+1} | NER损失: {losses['ner']:.4f}")

# 保存训练后的模型
nlp.to_disk("./gpa_ner_model")
  1. 注意事项:
    • 标注数据要覆盖尽可能多的GPA表达场景,提升模型泛化能力
    • 可以用Prodigy等工具批量标注简历数据,提高效率

内容的提问来源于stack exchange,提问作者Haneen Alahmadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 16:05:53