如何在Spacy中训练句子识别器?适配法律文本前置数字场景
训练Spacy分句器忽略法律案例前置数字
我需要从以数字开头的法律案例段落中提取句子,比如:
23 The main issue for determination therefore....
7 The Appellant’s factual witness...
目前Spacy的分句器会把整行(包括前置数字)识别为一个句子,我希望训练它忽略前置数字,从数字之后的内容开始标记句子起始。
我写了一段测试代码,但不确定是否正确,同时还有两个问题:
- 如何一次性传入多个训练示例
- 如何保存训练后的分句器并重复使用
我的测试代码:
import spacy from spacy.training import Example nlp = spacy.load('en_core_web_sm') senter = nlp.add_pipe("senter", name="my_senter") doc = nlp("32 It should be noted that each step") tokens_ref = ["32", "It", "should", "be", "noted", "that", "each", "step"] sents_ref = [False, True, False, False, False, False, False, False] examples = Example.from_dict(doc, {"words": tokens_ref, "sent_starts": sents_ref}) split_examples = examples.split_sents() senter.initialize(lambda: split_examples, nlp=nlp)
我是Spacy新手,相关训练资源较少,知道可以用其他方法移除前置数字,但希望通过训练Spacy实现该需求。
解决方案
先纠正现有代码的问题
你的代码仅完成了分句器的初始化,没有进行训练迭代更新参数,所以不会产生实际的效果。senter.initialize只是设置初始状态,必须通过训练循环调整模型参数才能让它学会你的规则。
1. 一次性传入多个训练示例
你可以准备一组训练数据,每条数据包含输入文本、对应的分词列表和句子起始标记,批量转换成Example对象即可。
示例训练数据准备:
train_data = [ # (输入文本, {"words": 分词列表, "sent_starts": 句子起始标记列表}) ("32 It should be noted that each step", { "words": ["32", "It", "should", "be", "noted", "that", "each", "step"], "sent_starts": [False, True, False, False, False, False, False, False] }), ("7 The Appellant’s factual witness provided key testimony.", { "words": ["7", "The", "Appellant’s", "factual", "witness", "provided", "key", "testimony", "."], "sent_starts": [False, True, False, False, False, False, False, False, False] }), ("23 The main issue for determination therefore requires further analysis.", { "words": ["23", "The", "main", "issue", "for", "determination", "therefore", "requires", "further", "analysis", "."], "sent_starts": [False, True, False, False, False, False, False, False, False, False, False] }) ]
批量生成训练示例:
examples = [] for text, annotations in train_data: doc = nlp.make_doc(text) # 用make_doc避免原模型的分句逻辑干扰标注 example = Example.from_dict(doc, annotations) examples.append(example)
2. 完整训练流程+保存复用模型
以下是包含训练循环、参数更新、模型保存与加载的完整代码:
import spacy from spacy.training import Example from spacy.util import minibatch, compounding # 加载基础模型,禁用自带的parser组件避免冲突 nlp = spacy.load('en_core_web_sm', disable=["parser"]) # 添加自定义分句器 senter = nlp.add_pipe("senter", name="custom_senter") # 准备训练数据(可根据实际场景补充更多样本) train_data = [ ("32 It should be noted that each step", { "words": ["32", "It", "should", "be", "noted", "that", "each", "step"], "sent_starts": [False, True, False, False, False, False, False, False] }), ("7 The Appellant’s factual witness provided key testimony.", { "words": ["7", "The", "Appellant’s", "factual", "witness", "provided", "key", "testimony", "."], "sent_starts": [False, True, False, False, False, False, False, False, False] }), ("23 The main issue for determination therefore requires further analysis.", { "words": ["23", "The", "main", "issue", "for", "determination", "therefore", "requires", "further", "analysis", "."], "sent_starts": [False, True, False, False, False, False, False, False, False, False, False] }) ] # 生成训练示例 examples = [] for text, annotations in train_data: doc = nlp.make_doc(text) example = Example.from_dict(doc, annotations) examples.append(example) # 初始化分句器 senter.initialize(lambda: examples, nlp=nlp) # 设置训练参数 optimizer = nlp.begin_training() # 批量大小随训练递增 batch_sizes = compounding(4.0, 32.0, 1.001) # 训练循环(可根据样本量调整epoch数量) for epoch in range(10): losses = {} batches = minibatch(examples, size=batch_sizes) for batch in batches: nlp.update(batch, sgd=optimizer, losses=losses) print(f"第{epoch+1}轮训练,损失值: {losses}") # 保存训练后的模型到本地 nlp.to_disk("./legal_senter_model") # 加载复用模型 loaded_nlp = spacy.load("./legal_senter_model") # 测试效果 test_texts = [ "32 It should be noted that each step", "7 The Appellant’s factual witness provided key testimony." ] for text in test_texts: doc = loaded_nlp(text) print(f"\n测试文本: {text}") print("识别出的句子:") for sent in doc.sents: print(f"- {sent.text}")
关键注意事项
- 加载基础模型时禁用
parser组件,因为它自带分句逻辑,会和我们的自定义senter冲突。 - 使用
nlp.make_doc生成原始文档,避免原模型的预处理逻辑干扰我们的训练标注。 - 训练样本要尽量覆盖实际场景中的各种情况(比如多位数、带点的数字如
23.等),样本越多模型效果越稳定。 - 训练轮数(
epoch)可根据样本数量调整,样本少则适当增加轮数,样本多则减少。
内容的提问来源于stack exchange,提问作者Vikneswaran Kumaran
相关产品推荐
相关产品推荐

