Kaggle NLP数据清洗时无法分词,文本出现粘连问题
问题分析与解决
你的代码导致单词粘连的核心原因是这一行:
text = re.sub('[^a-z]','',text) #removes non-alphabeticals
这个正则表达式会把所有非小写字母的字符(包括用于分隔单词的空格)全部替换为空,空格被删除后所有单词自然粘连成一串,后续的text.split()也就无法拆分出单个单词。
修正方案
将上述代码行修改为保留空格,仅移除非字母且非空格的字符:
text = re.sub('[^a-z\s]','',text) # 保留空格,仅移除非字母的非空格字符
另外,你代码里的text.replace('#', '')和text.replace('@', '')属于冗余操作——前面已经通过re.sub('[%s]' % re.escape(string.punctuation), '', text)移除了所有标点符号(包括#和@),这两行可以直接删除。
修正后的完整代码
import re import string from nltk.corpus import stopwords from nltk.stem import WordNetLemmatizer def text_cleaner(text): text = str(text).lower() # 转小写 text = re.sub('\d+', '', text) # 移除数字 text = re.sub('\[.*?\]','', text) # 移除方括号及内部内容 text = re.sub(r'https?://\S+|www\.\S+','',text) # 移除URL链接 text = re.sub(r'\bhtml\b', '', text) # 移除单独的"html"单词 # 移除表情符号 text = re.sub(r'[' u'\U0001F600-\U0001F64F' # 表情 u'\U0001F300-\U0001F5FF' # 符号与象形图 u'\U0001F680-\U0001F6FF' # 交通与地图符号 u'\U0001F1E0-\U0001F1FF' # 旗帜(iOS) u'\U00002702-\U000027B0' u'\U000024C2-\U0001F251' # 移除emoji ']+', '',text) text = re.sub('[%s]' % re.escape(string.punctuation), '', text) # 移除标点 text = re.sub('[^a-z\s]','',text) # 保留空格,移除非字母的非空格字符 text = stop_words(text) return text def stop_words(text): lem = WordNetLemmatizer() stop = set(stopwords.words('english')) stop.remove('not') punctuation = list(string.punctuation) stop.update(punctuation) text = text.split() text = [lem.lemmatize(word) for word in text if word not in stop] text = ' '.join(text) return text
修改后文本中的空格会被保留,stop_words函数里的text.split()可以正常拆分出单个单词,最终得到你期望的分词结果。
内容的提问来源于stack exchange,提问作者Francisco Vives
相关产品推荐
相关产品推荐

