You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Spacy中训练句子识别器?适配法律文本前置数字场景

训练Spacy分句器忽略法律案例前置数字

我需要从以数字开头的法律案例段落中提取句子,比如:

23 The main issue for determination therefore....
7 The Appellant’s factual witness...

目前Spacy的分句器会把整行(包括前置数字)识别为一个句子,我希望训练它忽略前置数字,从数字之后的内容开始标记句子起始。

我写了一段测试代码,但不确定是否正确,同时还有两个问题:

  1. 如何一次性传入多个训练示例
  2. 如何保存训练后的分句器并重复使用

我的测试代码:

import spacy
from spacy.training import Example

nlp = spacy.load('en_core_web_sm')
senter = nlp.add_pipe("senter", name="my_senter")

doc = nlp("32 It should be noted that each step")
tokens_ref = ["32", "It", "should", "be", "noted", "that", "each", "step"]
sents_ref = [False, True, False, False, False, False, False, False]
examples = Example.from_dict(doc, {"words": tokens_ref, "sent_starts": sents_ref})
split_examples = examples.split_sents()
senter.initialize(lambda: split_examples, nlp=nlp)

我是Spacy新手,相关训练资源较少,知道可以用其他方法移除前置数字,但希望通过训练Spacy实现该需求。


解决方案

先纠正现有代码的问题

你的代码仅完成了分句器的初始化,没有进行训练迭代更新参数,所以不会产生实际的效果。senter.initialize只是设置初始状态,必须通过训练循环调整模型参数才能让它学会你的规则。

1. 一次性传入多个训练示例

你可以准备一组训练数据,每条数据包含输入文本、对应的分词列表和句子起始标记,批量转换成Example对象即可。

示例训练数据准备:

train_data = [
    # (输入文本, {"words": 分词列表, "sent_starts": 句子起始标记列表})
    ("32 It should be noted that each step", {
        "words": ["32", "It", "should", "be", "noted", "that", "each", "step"],
        "sent_starts": [False, True, False, False, False, False, False, False]
    }),
    ("7 The Appellant’s factual witness provided key testimony.", {
        "words": ["7", "The", "Appellant’s", "factual", "witness", "provided", "key", "testimony", "."],
        "sent_starts": [False, True, False, False, False, False, False, False, False]
    }),
    ("23 The main issue for determination therefore requires further analysis.", {
        "words": ["23", "The", "main", "issue", "for", "determination", "therefore", "requires", "further", "analysis", "."],
        "sent_starts": [False, True, False, False, False, False, False, False, False, False, False]
    })
]

批量生成训练示例:

examples = []
for text, annotations in train_data:
    doc = nlp.make_doc(text)  # 用make_doc避免原模型的分句逻辑干扰标注
    example = Example.from_dict(doc, annotations)
    examples.append(example)

2. 完整训练流程+保存复用模型

以下是包含训练循环、参数更新、模型保存与加载的完整代码:

import spacy
from spacy.training import Example
from spacy.util import minibatch, compounding

# 加载基础模型,禁用自带的parser组件避免冲突
nlp = spacy.load('en_core_web_sm', disable=["parser"])
# 添加自定义分句器
senter = nlp.add_pipe("senter", name="custom_senter")

# 准备训练数据(可根据实际场景补充更多样本)
train_data = [
    ("32 It should be noted that each step", {
        "words": ["32", "It", "should", "be", "noted", "that", "each", "step"],
        "sent_starts": [False, True, False, False, False, False, False, False]
    }),
    ("7 The Appellant’s factual witness provided key testimony.", {
        "words": ["7", "The", "Appellant’s", "factual", "witness", "provided", "key", "testimony", "."],
        "sent_starts": [False, True, False, False, False, False, False, False, False]
    }),
    ("23 The main issue for determination therefore requires further analysis.", {
        "words": ["23", "The", "main", "issue", "for", "determination", "therefore", "requires", "further", "analysis", "."],
        "sent_starts": [False, True, False, False, False, False, False, False, False, False, False]
    })
]

# 生成训练示例
examples = []
for text, annotations in train_data:
    doc = nlp.make_doc(text)
    example = Example.from_dict(doc, annotations)
    examples.append(example)

# 初始化分句器
senter.initialize(lambda: examples, nlp=nlp)

# 设置训练参数
optimizer = nlp.begin_training()
# 批量大小随训练递增
batch_sizes = compounding(4.0, 32.0, 1.001)

# 训练循环(可根据样本量调整epoch数量)
for epoch in range(10):
    losses = {}
    batches = minibatch(examples, size=batch_sizes)
    for batch in batches:
        nlp.update(batch, sgd=optimizer, losses=losses)
    print(f"第{epoch+1}轮训练,损失值: {losses}")

# 保存训练后的模型到本地
nlp.to_disk("./legal_senter_model")

# 加载复用模型
loaded_nlp = spacy.load("./legal_senter_model")

# 测试效果
test_texts = [
    "32 It should be noted that each step",
    "7 The Appellant’s factual witness provided key testimony."
]
for text in test_texts:
    doc = loaded_nlp(text)
    print(f"\n测试文本: {text}")
    print("识别出的句子:")
    for sent in doc.sents:
        print(f"- {sent.text}")

关键注意事项

  • 加载基础模型时禁用parser组件,因为它自带分句逻辑,会和我们的自定义senter冲突。
  • 使用nlp.make_doc生成原始文档,避免原模型的预处理逻辑干扰我们的训练标注。
  • 训练样本要尽量覆盖实际场景中的各种情况(比如多位数、带点的数字如23.等),样本越多模型效果越稳定。
  • 训练轮数(epoch)可根据样本数量调整,样本少则适当增加轮数,样本多则减少。

内容的提问来源于stack exchange,提问作者Vikneswaran Kumaran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 05:30:49