如何通过TextVectorization函数解决分词连写错误问题?
解决TextVectorization中的连写单词分词问题
你的代码里存在两个核心问题导致连写单词出现:
- 最终返回未经过分词清洗的文本:你前面针对
words做了一系列清洗,但最后返回的processed_text是直接对原始小写文本lowercase做标点和数字移除的结果,前面的分词、停用词移除等操作完全没生效。 - 停用词处理方式有缺陷:循环替换停用词时用
' {i} '的格式,会导致停用词前后没有空格时无法被移除,反而可能造成单词拼接。
以下是修正后的custom_standardization函数:
def custom_standardization(input_data): import re import string from tensorflow import strings from nltk.corpus import stopwords stop_words = set(stopwords.words('english')) # 1. 文本转小写 lowercase = strings.lower(input_data) # 2. 移除HTML标签 stripped_html = strings.regex_replace(lowercase, "<br />", " ") # 3. 移除数字(含科学计数法格式) stripped_numbers = strings.regex_replace(stripped_html, r'\d+(?:\.\d*)?(?:[eE][+-]?\d+)?', ' ') # 4. 移除@提及内容 stripped_mentions = strings.regex_replace(stripped_numbers, r'@([A-Za-z0-9_]+)', ' ') # 5. 移除单个字符 stripped_single_chars = strings.regex_replace(stripped_mentions, r'\b\w\b', ' ') # 6. 移除标点符号 stripped_punct = strings.regex_replace(stripped_single_chars, '[%s]' % re.escape(string.punctuation), ' ') # 7. 批量移除停用词(避免循环,提升效率) stopword_pattern = r'\b(' + '|'.join(re.escape(word) for word in stop_words) + r')\b' stripped_stopwords = strings.regex_replace(stripped_punct, stopword_pattern, ' ') # 8. 清理多余空格(多空格转单空格,移除首尾空格) cleaned_text = strings.regex_replace(stripped_stopwords, r'\s+', ' ') cleaned_text = strings.strip(cleaned_text) return cleaned_text vectorize_layer = tf.keras.layers.TextVectorization( standardize=custom_standardization, max_tokens=vocab_size, output_mode='int', output_sequence_length=None, split='whitespace' # 此时split参数可正常生效 )
关键修改说明:
- 修正返回逻辑:所有清洗步骤连贯执行,最终返回完全处理后的文本,不再跳过中间的分词清洗操作
- 批量处理停用词:用正则表达式一次性匹配所有停用词,既提升效率又避免循环带来的匹配漏洞
- 清理多余空格:确保文本中仅用单个空格分隔单词,避免空字符串或多空格导致的分词异常
- 调整处理顺序:将标点移除放在停用词处理前,避免标点粘连导致停用词无法被正常匹配
修改后,TextVectorization的split='whitespace'参数可正常工作,不会再出现modifiedprospective、millionsyears这类连写单词。
内容的提问来源于stack exchange,提问作者Sreeharsha Kotta
相关产品推荐
相关产品推荐

