如何用PyTorch实现无标点全小写分词?求适配标点补全的PyTorch模型
无标点文本分词与标点补全的PyTorch实现方案
一、无标点、全小写文本的PyTorch分词实现
对于无标点、全小写的文本,直接使用Hugging Face Transformers库中的预训练tokenizer即可,这类tokenizer原生支持全小写文本处理,且能完美适配PyTorch环境。以下是示例代码:
from transformers import BertTokenizer import torch # 加载适配全小写文本的预训练tokenizer(bert-base-uncased专门针对小写文本训练) tokenizer = BertTokenizer.from_pretrained("bert-base-uncased") source_string = "first a modest refactor to fit the current project size and second a full refactor to move all our code into plugins dont feel like you have to code along to this whole book" # 执行分词,返回PyTorch张量格式的输入id、注意力掩码等 tokens = tokenizer( source_string, return_tensors="pt", # 指定返回PyTorch张量 truncation=True, padding="max_length", max_length=128 ) # 输出分词结果 print("分词后的输入ID:", tokens["input_ids"]) print("分词后的原始token:", tokenizer.convert_ids_to_tokens(tokens["input_ids"][0]))
说明:bert-base-uncased的tokenizer会自动处理全小写文本,无需额外转换;如果需要更轻量的模型,也可以选择distilbert-base-uncased,效果类似且推理速度更快。
二、标点补全的PyTorch模型选择与实现
针对英文标点补全场景,当spaCy和NLTK的轻量模型效果不佳时,推荐使用预训练序列标注类模型,这类模型在大规模文本语料上微调,能精准识别需要添加标点的位置,还能处理缩写形式(比如示例中的dont→don't)。
1. 推荐模型
- 多语言通用模型:
oliverguhr/fullstop-punctuation-multilingual(基于BERT训练,对英文标点补全表现出色) - 英文专用模型:
bert-base-uncased-finetuned-for-punctuation(训练数据聚焦英文,精度更优)
2. 实现示例
以oliverguhr/fullstop-punctuation-multilingual为例:
from transformers import pipeline # 加载标点补全pipeline,强制使用PyTorch作为后端 punctuator = pipeline( "token-classification", model="oliverguhr/fullstop-punctuation-multilingual", tokenizer="oliverguhr/fullstop-punctuation-multilingual", framework="pt" ) source_string = "first a modest refactor to fit the current project size and second a full refactor to move all our code into plugins dont feel like you have to code along to this whole book" # 执行标点补全 results = punctuator(source_string) # 拼接结果,生成带标点的文本 punctuated_text = "" prev_token = "" for item in results: token = item["word"] # 处理分词器的子词拆分(比如"don't"可能被拆为"don"和"##t") if token.startswith("##"): token = token[2:] punctuated_text += token else: if prev_token: punctuated_text += " " punctuated_text += token # 添加对应的标点(模型返回的entity为标点符号或"0"表示无标点) if item["entity"] != "0": punctuated_text += item["entity"] # 输出补全后的文本 print("补全标点后的文本:", punctuated_text)
3. 自定义微调建议
如果现有预训练模型无法满足特定场景需求,可以收集带标点/无标点的平行语料,将任务转为序列标注任务(每个token对应是否添加标点的标签),用PyTorch搭建BERT/LSTM等模型进行自定义微调,进一步提升精度。
内容的提问来源于stack exchange,提问作者Daniel
相关产品推荐
相关产品推荐

