Python中如何将指定多词短语合并为单个token进行分词?
Hey there!看起来你需要把像'red blood cell'、'platelet count'这类多词短语识别成单个token,而不是拆成独立单词。下面是几个实用的实现方案,你可以根据自己的场景选择:
1. 基于预定义词典的匹配法
这是最直接的方案——先维护一个包含所有目标多词短语的词典,然后在分词时优先匹配这些短语,避免被拆分成单个单词。核心思路是先匹配长短语,再匹配短短语,防止短短语提前匹配导致长短语被拆分(比如先匹配"blood cell"就会破坏"red blood cell"的完整识别)。
示例代码(Python):
import re def tokenize_with_phrases(text, phrase_list): # 按短语的单词数量倒序排序,确保长短语优先匹配 sorted_phrases = sorted(phrase_list, key=lambda x: len(x.split()), reverse=True) # 构建正则匹配模式,转义短语中的特殊字符,避免正则语法冲突 pattern = re.compile(r'\b(' + '|'.join(re.escape(p) for p in sorted_phrases) + r')\b', re.IGNORECASE) # 用特殊占位符替换匹配到的短语,方便后续分割 placeholder = "__PHRASE__" replaced_text = pattern.sub(lambda m: f"{placeholder}{m.group(0)}{placeholder}", text) tokens = [] for part in replaced_text.split(): if part.startswith(placeholder) and part.endswith(placeholder): # 提取占位符包裹的短语,转为小写并清理 clean_phrase = part.strip(placeholder).lower() tokens.append(clean_phrase) else: # 处理单个单词,去除标点符号 clean_word = re.sub(r'[^\w\s]', '', part.lower()) if clean_word: tokens.append(clean_word) return tokens # 测试用例 input_str = "hello my name is vishal, can you please help me with the red blood cells and platelet count. The white blood cell is a single word." target_phrases = ["red blood cell", "platelet count", "white blood cell"] result = tokenize_with_phrases(input_str, target_phrases) print(result)
这个方案的优点是实现简单、速度快,适合短语列表明确且变动不大的场景;缺点是需要手动维护词典,无法识别未提前定义的短语。
2. 基于词性规则的短语识别
如果你的目标短语有固定的语法结构(比如医学术语常是「形容词+名词+名词」或「名词+名词」),可以利用**词性标注(POS Tagging)**来自动识别符合规则的多词组合。
示例代码(用spaCy实现):
import spacy import re # 加载spaCy的英文基础模型(需先安装:pip install spacy && python -m spacy download en_core_web_sm) nlp = spacy.load("en_core_web_sm") def tokenize_with_pos_rules(text): doc = nlp(text) tokens = [] i = 0 doc_len = len(doc) while i < doc_len: # 匹配「形容词+名词+名词」结构(比如red blood cell) if i + 2 < doc_len and doc[i].pos_ == "ADJ" and doc[i+1].pos_ == "NOUN" and doc[i+2].pos_ == "NOUN": phrase = f"{doc[i].text.lower()} {doc[i+1].text.lower()} {doc[i+2].text.lower()}" tokens.append(phrase) i += 3 # 匹配「名词+名词」结构(比如platelet count) elif i + 1 < doc_len and doc[i].pos_ == "NOUN" and doc[i+1].pos_ == "NOUN": phrase = f"{doc[i].text.lower()} {doc[i+1].text.lower()}" tokens.append(phrase) i += 2 else: # 处理单个单词,清理标点 clean_word = re.sub(r'[^\w\s]', '', doc[i].text.lower()) if clean_word: tokens.append(clean_word) i += 1 return tokens # 测试用例 result = tokenize_with_pos_rules(input_str) print(result)
这个方案的优点是不需要手动维护短语列表,能自动识别符合规则的短语;缺点是规则需要根据领域调整,对无固定结构的短语识别效果差。
3. 基于预训练领域模型的实体/短语抽取
如果处理的是专业领域文本(比如你的例子是医学文本),可以用预训练的领域语言模型来自动识别专业术语,这类模型已经在大量领域语料上训练过,能准确识别多词短语。
示例代码(用医学领域的spaCy模型):
import spacy import re # 安装医学领域模型:pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_sm-0.5.4.tar.gz nlp = spacy.load("en_core_sci_sm") def tokenize_with_domain_model(text): doc = nlp(text) tokens = [] ent_start_idx = 0 # 先处理模型识别出的实体(专业术语) for ent in doc.ents: # 添加实体之前的单个单词 for token in doc[ent_start_idx:ent.start]: clean_word = re.sub(r'[^\w\s]', '', token.text.lower()) if clean_word: tokens.append(clean_word) # 将实体作为单个token加入列表 tokens.append(ent.text.lower()) ent_start_idx = ent.end # 添加剩余的单个单词 for token in doc[ent_start_idx:]: clean_word = re.sub(r'[^\w\s]', '', token.text.lower()) if clean_word: tokens.append(clean_word) return tokens # 测试用例 result = tokenize_with_domain_model(input_str) print(result)
这个方案的优点是准确率高,能自动识别未提前定义的专业短语;缺点是需要加载较大的模型,对非领域文本效果一般。
4. 训练自定义分词器
如果需要长期处理大量领域文本,且希望分词器能自动学习常见的多词短语,可以训练自定义分词器(比如基于BPE算法的分词器)。
示例代码(用Hugging Face Tokenizers):
from tokenizers import ByteLevelBPETokenizer import re # 准备训练语料:可以是包含目标短语的文本文件,或者直接用短语列表+输入文本 phrase_list = ["red blood cell", "platelet count", "white blood cell"] with open("domain_corpus.txt", "w", encoding="utf-8") as f: f.write("\n".join(phrase_list) + "\n" + input_str) # 初始化BPE分词器并训练 tokenizer = ByteLevelBPETokenizer() tokenizer.train(files=["domain_corpus.txt"], vocab_size=1000, min_frequency=1) # 对输入文本进行分词 encoding = tokenizer.encode(input_str) # 将token ID转换为文本 raw_tokens = [tokenizer.decode([token_id]) for token_id in encoding.ids] # 清理token,去除标点和多余空格 clean_tokens = [] for token in raw_tokens: cleaned = re.sub(r'[^\w\s]', '', token.lower()).strip() if cleaned: clean_tokens.append(cleaned) print(clean_tokens)
这个方案的优点是能自适应领域文本,长期使用效果会越来越好;缺点是需要准备训练语料,初期配置成本较高。
内容的提问来源于stack exchange,提问作者VISHAL GADHVI

