如何使用Gensim的remove_stopwords批量移除pandas列中的停用词?
实现方法
首先导入所需的依赖库:
import pandas as pd from gensim.parsing.preprocessing import remove_stopwords
方法1:使用pandas apply方法(最推荐,无需手动写循环)
pandas的apply方法会自动遍历列中的每一个元素执行指定函数,比手动写循环效率更高、代码更简洁:
# 构造示例数据 data = {'Text':['I like it.', 'He is happy.', 'This is apple.', 'It is great.']} df = pd.DataFrame(data) # 直接对Text列调用apply方法传入停用词移除函数,结果存入新列 df['text_without_stopwords'] = df['Text'].apply(remove_stopwords)
方法2:手动写循环实现
如果你需要用显式循环完成处理,可以逐行遍历DataFrame赋值:
# 构造示例数据 data = {'Text':['I like it.', 'He is happy.', 'This is apple.', 'It is great.']} df = pd.DataFrame(data) # 初始化新列 df['text_without_stopwords'] = '' # 遍历每一行索引 for idx in df.index: # 取出当前行的Text内容处理后,赋值给新列对应位置 df.loc[idx, 'text_without_stopwords'] = remove_stopwords(df.loc[idx, 'Text'])
两种方法运行后得到的结果一致,新列的处理结果如下:
| 索引 | Text | text_without_stopwords |
|---|---|---|
| 0 | I like it. | like. |
| 1 | He is happy. | happy. |
| 2 | This is apple. | apple. |
| 3 | It is great. | great. |
内容的提问来源于stack exchange,提问作者MC2020
相关产品推荐
相关产品推荐

