You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清空Pandas DataFrame中不含指定词汇的单元格并添加分类标签

Pandas分词文本内容过滤与标签标注实现方案

实现思路

  • 提前将正负例天气词汇存为集合,相比列表可大幅提升匹配效率
  • 编写通用的单单元格处理函数,一次性完成内容过滤和标签判断逻辑,避免重复编码
  • 批量遍历所有text开头的文本列,自动生成对应处理后的内容列和标签列,全程保留id列数据不变

完整可运行代码

import pandas as pd
import numpy as np

# 1. 定义正负例词汇
pos = {'heat', 'sun'}
neg = {'cold', 'rain'}
all_weather_words = pos.union(neg)

# 2. 单元格处理函数:输入分词列表,返回(处理后内容, 标签)
def process_cell(cell):
    # 处理空值、非列表、空列表的异常情况
    if pd.isna(cell) or not isinstance(cell, list) or len(cell) == 0:
        return np.nan, ''
    # 无天气相关词汇直接清空
    if not any(word in all_weather_words for word in cell):
        return np.nan, ''
    # 匹配标签:同时存在正负例时默认优先正例,可按需调整逻辑
    if any(word in pos for word in cell):
        return cell, 'pos'
    elif any(word in neg for word in cell):
        return cell, 'neg'
    return np.nan, ''

# 3. 构造测试数据(你自己使用时替换为自己的DataFrame即可)
data = {
    'id': [123, 124, 125, 126],
    'text1': [['it', 'was', 'cold'], np.nan, ['the', 'heat'], np.nan],
    'text2': [["i", "wasn't", 'there'], ['hello', 'there'], ['the', 'cold'], ['the', 'heat']]
}
df = pd.DataFrame(data)

# 4. 批量处理所有文本列
text_columns = [col for col in df.columns if col.startswith('text')]
for idx, col in enumerate(text_columns, 1):
    df[[col, f'label{idx}']] = df[col].apply(lambda x: pd.Series(process_cell(x)))

# 可选:将空值替换为空字符串,和示例输出格式完全一致
df = df.fillna('')

# 查看结果
print(df)

注意事项

  • 若你的DataFrame里的分词列表是字符串格式(比如存储时转成了"['it', 'was', 'cold']"这类字符串),需要先用ast.literal_eval将字符串转换为Python列表再进行处理
  • 若单元格同时存在正负例词汇,可自行调整process_cell里的标签判断逻辑,比如返回both这类混合标签

内容的提问来源于stack exchange,提问作者Leonie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 10:45:01