You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从150k行40列数据集移除含特定关键词的单元格

嘿,我来帮你搞定这个大型数据集的处理需求!针对你的150k行、40列混合类型数据集,我们可以用Python的pandas库高效完成「移除含指定关键词的单元格并替换为NA」的任务,下面是具体方案:

处理大型数据集:移除含指定关键词的单元格并替换为NA

核心思路

我们会遍历数据集中的每一个单元格,判断其是否包含目标关键词:如果是字符串/字符类型且包含关键词,就替换为NaN(对应你示例里的NA);数字类型或不包含关键词的内容则保持原样。这种方法既能兼容混合数据类型,又能高效处理大体积数据集。

具体实现步骤

1. 准备工具

先确保你安装了pandas(处理表格数据的神器),如果没装,运行以下命令:

pip install pandas

2. 核心代码

import pandas as pd
import numpy as np

# 读取你的数据集(如果是Excel格式,改成pd.read_excel,其他格式按需调整)
df = pd.read_csv("your_dataset.csv")

# 定义要匹配的关键词,比如你示例里的"is"
target_keyword = "is"

# 定义单元格处理函数
def check_and_replace(cell):
    # 跳过空值,避免报错
    if pd.isna(cell):
        return cell
    # 只处理字符串类型单元格,数字类型直接返回
    if isinstance(cell, str):
        # 如果包含关键词,返回NaN(即示例中的NA)
        if target_keyword in cell:
            return np.nan
    # 不符合替换条件的内容直接返回
    return cell

# 把函数应用到整个数据集,逐单元格处理
df_processed = df.applymap(check_and_replace)

# 保存处理后的数据集
df_processed.to_csv("processed_dataset.csv", index=False)

实用扩展说明

  • 忽略大小写匹配:如果需要不区分大小写(比如匹配"Is""IS"),把判断条件改成:
    if target_keyword.lower() in str(cell).lower():
    
  • 多关键词匹配:如果要同时移除含多个关键词的单元格,比如["is", "are"],可以修改函数:
    target_keywords = {"is", "are"}
    def check_and_replace(cell):
        if pd.isna(cell):
            return cell
        if isinstance(cell, str):
            if any(keyword in cell for keyword in target_keywords):
                return np.nan
        return cell
    
  • 效率升级:如果你的数据集特别庞大,还可以用swifter库加速处理(先装pip install swifter,然后把df.applymap(...)改成df.swifter.applymap(...))。

示例验证

用你给出的测试数据跑一遍,完全符合预期:

初始数据集:

AB
1My name is Sam.
Hello2
Who are youThe water is green.

处理后(关键词为"is"):

AB
1NaN
Hello2
Who are youNaN

内容的提问来源于stack exchange,提问作者Gothram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:58:14