如何用Python将上标与词根分离并实现独立分词?
如何将上标/非ASCII符号与词根分离为独立分词Token
方案1:预处理文本插入空格,再分词
核心思路是在非ASCII符号(或指定的上标字符)前插入空格,让分词工具自动将其识别为独立token。
针对所有非ASCII字符的实现
import re import nltk # 可替换为你使用的分词库(如spaCy、jieba等) text = "This is a sentence about testString™" # 在每个非ASCII字符前插入空格 processed_text = re.sub(r'([^\x00-\x7F])', r' \1', text) # 处理后文本:"This is a sentence about testString ™" # 执行分词 tokens = nltk.word_tokenize(processed_text.lower()) print(tokens) # 输出: ['this', 'is', 'a', 'sentence', 'about', 'teststring', '™']
仅针对上标字符的精准实现
如果只想处理上标(而非所有非ASCII字符),可以匹配上标的Unicode专属范围:
import re import nltk text = "This is a sentence about testString™ x² y³" # 匹配Unicode上标字符(包含上标数字、字母等) superscript_pattern = r'([\u00B2\u00B3\u00B9\u2070-\u207E\u207F])' processed_text = re.sub(superscript_pattern, r' \1', text) # 处理后文本:"This is a sentence about testString ™ x ² y ³" tokens = nltk.word_tokenize(processed_text.lower()) print(tokens) # 输出: ['this', 'is', 'a', 'sentence', 'about', 'teststring', '™', 'x', '²', 'y', '³']
方案2:后处理分词结果,拆分混合token
如果已经有了初始分词结果,可以直接遍历每个token,拆分出其中的ASCII词根和非ASCII符号:
def split_mixed_token(token): ascii_chars = [] non_ascii_chars = [] for char in token: if ord(char) <= 127: ascii_chars.append(char) else: non_ascii_chars.append(char) result = [] if ascii_chars: result.append(''.join(ascii_chars)) if non_ascii_chars: result.append(''.join(non_ascii_chars)) return result original_tokens = ["this", "is", "a", "sentence", "about", "testString™"] final_tokens = [] for token in original_tokens: final_tokens.extend(split_mixed_token(token)) print(final_tokens) # 输出: ['this', 'is', 'a', 'sentence', 'about', 'testString', '™']
思路说明
- 方案1适合在分词前预处理文本,兼容性好,大部分分词工具都能识别空格分隔的独立单元。
- 方案2适合已有固定分词流程的场景,无需修改文本预处理逻辑,直接调整分词结果即可。
内容的提问来源于stack exchange,提问作者alex
相关产品推荐
相关产品推荐

