You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:无依赖库实现符合特殊规则的文本分词(Tokenizing)函数

自定义规则文本分词实现(无第三方库依赖)

核心思路

用Python内置的re模块,通过正则表达式正向匹配(而非拆分)精准捕获所有符合规则的token,避免split()方法误拆需要保留内容的问题。先处理收缩词的展开,再用正则匹配所有需保留的整体和单独拆分的标点。

实现代码

import re

# 常见收缩词映射,用于展开如I’m→I am
CONTRACTIONS = {
    "I’m": "I am",
    "don’t": "do not",
    "can’t": "cannot",
    "won’t": "will not",
    "it’s": "it is",
    "you’re": "you are",
    "we’re": "we are",
    "they’re": "they are",
    "isn’t": "is not",
    "aren’t": "are not",
    "wasn’t": "was not",
    "weren’t": "were not",
    "hasn’t": "has not",
    "haven’t": "have not",
    "hadn’t": "had not",
    "doesn’t": "does not",
    "didn’t": "did not",
    "wouldn’t": "would not",
    "couldn’t": "could not",
    "shouldn’t": "should not",
    "that’s": "that is",
    "there’s": "there is"
}

def replace_contractions(text):
    """替换文本中的收缩词为展开形式"""
    for contraction, expansion in CONTRACTIONS.items():
        text = text.replace(contraction, expansion)
    return text

def tokenize(text):
    # 第一步:处理收缩词展开
    processed_text = replace_contractions(text)
    
    # 正则匹配模式:按优先级匹配需要保留的整体,再匹配单词和单独标点
    token_pattern = r'''
        \d{1,3}(?:,\d{3})+|          # 带千位逗号的数字(如332,087,410)
        \d+\.\d+|                    # 带小数点的数字(如3.14)
        (?:[A-Za-z]\.)+|             # 缩写/首字母缩略词(如U.S.A.、Mr.)
        \d{1,2}/\d{1,2}/\d{2,4}|     # 日期格式(如01/22/2021)
        \w+(?:-\w+)+|                # 连字符连接的短语(如state-of-the-art)
        \w+|                         # 普通单词(包括所有格前的部分)
        [’,.;!?]                     # 需要单独拆分的标点(逗号、句点、撇号等)
    '''
    
    # 匹配所有token并过滤空内容
    tokens = re.findall(token_pattern, processed_text, re.VERBOSE)
    tokens = [token.strip() for token in tokens if token.strip()]
    
    return tokens

# 测试示例
if __name__ == "__main__":
    sample_input = "The current population of U.S.A. is 332,087,410 as of Friday, 01/22/2021, based on Worldometer elaboration of the latest United Nations’ data."
    result = tokenize(sample_input)
    print(' '.join(result))

规则对应说明

  1. 句点处理:正则中(?:[A-Za-z]\.)+匹配U.S.A.这类缩略词,\d+\.\d+匹配带小数点的数字,确保这类句点不被拆分;单独的句点会被匹配为独立token。
  2. 撇号处理:先通过replace_contractions展开收缩词(如I’m→I am);所有格的撇号(如United Nations’)会被正则匹配为独立的’token,实现拆分。
  3. 日期保留:\d{1,2}/\d{1,2}/\d{2,4}匹配MM/DD/YYYY格式的日期,确保整体不拆分。
  4. 连字符短语:\w+(?:-\w+)+匹配连字符连接的复合短语,保持整体。
  5. 数字中的逗号:\d{1,3}(?:,\d{3})+匹配带千位逗号的数字,确保逗号不被拆分;普通位置的逗号会被匹配为独立token。

测试输出

运行代码后,示例输入的输出与要求一致:

The current population of U.S.A. is 332,087,410 as of Friday , 01/22/2021 , based on Worldometer elaboration of the latest United Nations ’ data .

内容的提问来源于stack exchange,提问作者Jericho Sims

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 17:25:19