You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PyTorch实现无标点全小写分词?求适配标点补全的PyTorch模型

无标点文本分词与标点补全的PyTorch实现方案

一、无标点、全小写文本的PyTorch分词实现

对于无标点、全小写的文本,直接使用Hugging Face Transformers库中的预训练tokenizer即可,这类tokenizer原生支持全小写文本处理,且能完美适配PyTorch环境。以下是示例代码:

from transformers import BertTokenizer
import torch

# 加载适配全小写文本的预训练tokenizer(bert-base-uncased专门针对小写文本训练)
tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")

source_string = "first a modest refactor to fit the current project size and second a full refactor to move all our code into plugins dont feel like you have to code along to this whole book"

# 执行分词,返回PyTorch张量格式的输入id、注意力掩码等
tokens = tokenizer(
    source_string,
    return_tensors="pt",  # 指定返回PyTorch张量
    truncation=True,
    padding="max_length",
    max_length=128
)

# 输出分词结果
print("分词后的输入ID:", tokens["input_ids"])
print("分词后的原始token:", tokenizer.convert_ids_to_tokens(tokens["input_ids"][0]))

说明:bert-base-uncased的tokenizer会自动处理全小写文本,无需额外转换;如果需要更轻量的模型,也可以选择distilbert-base-uncased,效果类似且推理速度更快。

二、标点补全的PyTorch模型选择与实现

针对英文标点补全场景,当spaCy和NLTK的轻量模型效果不佳时,推荐使用预训练序列标注类模型,这类模型在大规模文本语料上微调,能精准识别需要添加标点的位置,还能处理缩写形式(比如示例中的dont→don't)。

1. 推荐模型

  • 多语言通用模型:oliverguhr/fullstop-punctuation-multilingual(基于BERT训练,对英文标点补全表现出色)
  • 英文专用模型:bert-base-uncased-finetuned-for-punctuation(训练数据聚焦英文,精度更优)

2. 实现示例

以oliverguhr/fullstop-punctuation-multilingual为例:

from transformers import pipeline

# 加载标点补全pipeline,强制使用PyTorch作为后端
punctuator = pipeline(
    "token-classification",
    model="oliverguhr/fullstop-punctuation-multilingual",
    tokenizer="oliverguhr/fullstop-punctuation-multilingual",
    framework="pt"
)

source_string = "first a modest refactor to fit the current project size and second a full refactor to move all our code into plugins dont feel like you have to code along to this whole book"

# 执行标点补全
results = punctuator(source_string)

# 拼接结果,生成带标点的文本
punctuated_text = ""
prev_token = ""
for item in results:
    token = item["word"]
    # 处理分词器的子词拆分(比如"don't"可能被拆为"don"和"##t")
    if token.startswith("##"):
        token = token[2:]
        punctuated_text += token
    else:
        if prev_token:
            punctuated_text += " "
        punctuated_text += token
    # 添加对应的标点(模型返回的entity为标点符号或"0"表示无标点)
    if item["entity"] != "0":
        punctuated_text += item["entity"]

# 输出补全后的文本
print("补全标点后的文本:", punctuated_text)

3. 自定义微调建议

如果现有预训练模型无法满足特定场景需求,可以收集带标点/无标点的平行语料,将任务转为序列标注任务(每个token对应是否添加标点的标签),用PyTorch搭建BERT/LSTM等模型进行自定义微调,进一步提升精度。

内容的提问来源于stack exchange,提问作者Daniel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 15:17:52