如何清空Pandas DataFrame中不含指定词汇的单元格并添加分类标签
Pandas分词文本内容过滤与标签标注实现方案
实现思路
- 提前将正负例天气词汇存为集合,相比列表可大幅提升匹配效率
- 编写通用的单单元格处理函数,一次性完成内容过滤和标签判断逻辑,避免重复编码
- 批量遍历所有
text开头的文本列,自动生成对应处理后的内容列和标签列,全程保留id列数据不变
完整可运行代码
import pandas as pd import numpy as np # 1. 定义正负例词汇 pos = {'heat', 'sun'} neg = {'cold', 'rain'} all_weather_words = pos.union(neg) # 2. 单元格处理函数:输入分词列表,返回(处理后内容, 标签) def process_cell(cell): # 处理空值、非列表、空列表的异常情况 if pd.isna(cell) or not isinstance(cell, list) or len(cell) == 0: return np.nan, '' # 无天气相关词汇直接清空 if not any(word in all_weather_words for word in cell): return np.nan, '' # 匹配标签:同时存在正负例时默认优先正例,可按需调整逻辑 if any(word in pos for word in cell): return cell, 'pos' elif any(word in neg for word in cell): return cell, 'neg' return np.nan, '' # 3. 构造测试数据(你自己使用时替换为自己的DataFrame即可) data = { 'id': [123, 124, 125, 126], 'text1': [['it', 'was', 'cold'], np.nan, ['the', 'heat'], np.nan], 'text2': [["i", "wasn't", 'there'], ['hello', 'there'], ['the', 'cold'], ['the', 'heat']] } df = pd.DataFrame(data) # 4. 批量处理所有文本列 text_columns = [col for col in df.columns if col.startswith('text')] for idx, col in enumerate(text_columns, 1): df[[col, f'label{idx}']] = df[col].apply(lambda x: pd.Series(process_cell(x))) # 可选:将空值替换为空字符串,和示例输出格式完全一致 df = df.fillna('') # 查看结果 print(df)
注意事项
- 若你的DataFrame里的分词列表是字符串格式(比如存储时转成了
"['it', 'was', 'cold']"这类字符串),需要先用ast.literal_eval将字符串转换为Python列表再进行处理 - 若单元格同时存在正负例词汇,可自行调整
process_cell里的标签判断逻辑,比如返回both这类混合标签
内容的提问来源于stack exchange,提问作者Leonie
相关产品推荐
相关产品推荐

