You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将指定多词短语合并为单个token进行分词?

Hey there!看起来你需要把像'red blood cell'、'platelet count'这类多词短语识别成单个token,而不是拆成独立单词。下面是几个实用的实现方案,你可以根据自己的场景选择:

1. 基于预定义词典的匹配法

这是最直接的方案——先维护一个包含所有目标多词短语的词典,然后在分词时优先匹配这些短语,避免被拆分成单个单词。核心思路是先匹配长短语,再匹配短短语,防止短短语提前匹配导致长短语被拆分(比如先匹配"blood cell"就会破坏"red blood cell"的完整识别)。

示例代码(Python):

import re

def tokenize_with_phrases(text, phrase_list):
    # 按短语的单词数量倒序排序,确保长短语优先匹配
    sorted_phrases = sorted(phrase_list, key=lambda x: len(x.split()), reverse=True)
    # 构建正则匹配模式,转义短语中的特殊字符,避免正则语法冲突
    pattern = re.compile(r'\b(' + '|'.join(re.escape(p) for p in sorted_phrases) + r')\b', re.IGNORECASE)
    
    # 用特殊占位符替换匹配到的短语,方便后续分割
    placeholder = "__PHRASE__"
    replaced_text = pattern.sub(lambda m: f"{placeholder}{m.group(0)}{placeholder}", text)
    
    tokens = []
    for part in replaced_text.split():
        if part.startswith(placeholder) and part.endswith(placeholder):
            # 提取占位符包裹的短语,转为小写并清理
            clean_phrase = part.strip(placeholder).lower()
            tokens.append(clean_phrase)
        else:
            # 处理单个单词,去除标点符号
            clean_word = re.sub(r'[^\w\s]', '', part.lower())
            if clean_word:
                tokens.append(clean_word)
    return tokens

# 测试用例
input_str = "hello my name is vishal, can you please help me with the red blood cells and platelet count. The white blood cell is a single word."
target_phrases = ["red blood cell", "platelet count", "white blood cell"]
result = tokenize_with_phrases(input_str, target_phrases)
print(result)

这个方案的优点是实现简单、速度快,适合短语列表明确且变动不大的场景;缺点是需要手动维护词典,无法识别未提前定义的短语。

2. 基于词性规则的短语识别

如果你的目标短语有固定的语法结构(比如医学术语常是「形容词+名词+名词」或「名词+名词」),可以利用**词性标注(POS Tagging)**来自动识别符合规则的多词组合。

示例代码(用spaCy实现):

import spacy
import re

# 加载spaCy的英文基础模型(需先安装:pip install spacy && python -m spacy download en_core_web_sm)
nlp = spacy.load("en_core_web_sm")

def tokenize_with_pos_rules(text):
    doc = nlp(text)
    tokens = []
    i = 0
    doc_len = len(doc)
    
    while i < doc_len:
        # 匹配「形容词+名词+名词」结构(比如red blood cell)
        if i + 2 < doc_len and doc[i].pos_ == "ADJ" and doc[i+1].pos_ == "NOUN" and doc[i+2].pos_ == "NOUN":
            phrase = f"{doc[i].text.lower()} {doc[i+1].text.lower()} {doc[i+2].text.lower()}"
            tokens.append(phrase)
            i += 3
        # 匹配「名词+名词」结构(比如platelet count)
        elif i + 1 < doc_len and doc[i].pos_ == "NOUN" and doc[i+1].pos_ == "NOUN":
            phrase = f"{doc[i].text.lower()} {doc[i+1].text.lower()}"
            tokens.append(phrase)
            i += 2
        else:
            # 处理单个单词,清理标点
            clean_word = re.sub(r'[^\w\s]', '', doc[i].text.lower())
            if clean_word:
                tokens.append(clean_word)
            i += 1
    return tokens

# 测试用例
result = tokenize_with_pos_rules(input_str)
print(result)

这个方案的优点是不需要手动维护短语列表,能自动识别符合规则的短语;缺点是规则需要根据领域调整,对无固定结构的短语识别效果差。

3. 基于预训练领域模型的实体/短语抽取

如果处理的是专业领域文本(比如你的例子是医学文本),可以用预训练的领域语言模型来自动识别专业术语,这类模型已经在大量领域语料上训练过,能准确识别多词短语。

示例代码(用医学领域的spaCy模型):

import spacy
import re

# 安装医学领域模型:pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_sm-0.5.4.tar.gz
nlp = spacy.load("en_core_sci_sm")

def tokenize_with_domain_model(text):
    doc = nlp(text)
    tokens = []
    ent_start_idx = 0
    
    # 先处理模型识别出的实体(专业术语)
    for ent in doc.ents:
        # 添加实体之前的单个单词
        for token in doc[ent_start_idx:ent.start]:
            clean_word = re.sub(r'[^\w\s]', '', token.text.lower())
            if clean_word:
                tokens.append(clean_word)
        # 将实体作为单个token加入列表
        tokens.append(ent.text.lower())
        ent_start_idx = ent.end
    
    # 添加剩余的单个单词
    for token in doc[ent_start_idx:]:
        clean_word = re.sub(r'[^\w\s]', '', token.text.lower())
        if clean_word:
            tokens.append(clean_word)
    return tokens

# 测试用例
result = tokenize_with_domain_model(input_str)
print(result)

这个方案的优点是准确率高,能自动识别未提前定义的专业短语;缺点是需要加载较大的模型,对非领域文本效果一般。

4. 训练自定义分词器

如果需要长期处理大量领域文本,且希望分词器能自动学习常见的多词短语,可以训练自定义分词器(比如基于BPE算法的分词器)。

示例代码(用Hugging Face Tokenizers):

from tokenizers import ByteLevelBPETokenizer
import re

# 准备训练语料:可以是包含目标短语的文本文件,或者直接用短语列表+输入文本
phrase_list = ["red blood cell", "platelet count", "white blood cell"]
with open("domain_corpus.txt", "w", encoding="utf-8") as f:
    f.write("\n".join(phrase_list) + "\n" + input_str)

# 初始化BPE分词器并训练
tokenizer = ByteLevelBPETokenizer()
tokenizer.train(files=["domain_corpus.txt"], vocab_size=1000, min_frequency=1)

# 对输入文本进行分词
encoding = tokenizer.encode(input_str)
# 将token ID转换为文本
raw_tokens = [tokenizer.decode([token_id]) for token_id in encoding.ids]

# 清理token,去除标点和多余空格
clean_tokens = []
for token in raw_tokens:
    cleaned = re.sub(r'[^\w\s]', '', token.lower()).strip()
    if cleaned:
        clean_tokens.append(cleaned)
print(clean_tokens)

这个方案的优点是能自适应领域文本,长期使用效果会越来越好;缺点是需要准备训练语料,初期配置成本较高。


内容的提问来源于stack exchange,提问作者VISHAL GADHVI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:10:39