You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame指定列按条件批量修改值的实现方法

嘿,我来帮你搞定这个批量处理DataFrame列的需求!不用逐个指定列名就能实现规则映射,下面给你分场景讲解具体实现:

批量实现DataFrame列值映射:>=10→100,<10→0

首先咱们先把你给出的初始化代码补全,生成一个可测试的示例DataFrame:

import numpy as np
import pandas as pd
from sklearn.datasets import make_classification

# 生成带6个特征的示例数据集
X, y = make_classification(n_samples=1000, n_features=6, n_informative=3, n_classes=2, random_state=42)
df = pd.DataFrame(X, columns=[f"col_{i}" for i in range(6)])

场景1:处理所有数值型列

如果你的需求是对DataFrame里所有数值类型的列统一应用规则,完全不用手动列名,用下面两种方法都能搞定:

方法1:用np.where(效率优先,适合大数据集)

这是向量级别的操作,比循环快很多:

# 自动筛选所有数值型列
numeric_cols = df.select_dtypes(include=[np.number]).columns
# 批量应用映射规则
df[numeric_cols] = np.where(df[numeric_cols] >= 10, 100, 0)

方法2:用applymap(代码更直观,易读)

自定义一个映射函数,再批量应用到目标列:

def map_value(x):
    return 100 if x >= 10 else 0

# 筛选数值列后应用函数
numeric_cols = df.select_dtypes(include=[np.number]).columns
df[numeric_cols] = df[numeric_cols].applymap(map_value)

场景2:处理自定义批量列(不用逐个写列名)

如果你只需要处理特定的一批列(比如列名有规律,或者指定范围),可以通过列名筛选来批量获取目标列:

# 示例1:处理前4列
target_cols = df.columns[:4]

# 示例2:处理列名以"col_"开头的列(适合有命名规则的列)
target_cols = [col for col in df.columns if col.startswith("col_")]

# 示例3:处理索引为偶数的列
target_cols = df.columns[::2]

# 最后统一应用规则
df[target_cols] = np.where(df[target_cols] >= 10, 100, 0)

验证结果是否符合预期

可以用下面的代码快速检查映射是否生效:

# 查看前5行数据
print(df.head())
# 统计被替换为100的单元格数量
print("替换为100的单元格总数:", df[df == 100].count().sum())

这样不管是全量数值列还是自定义批量列,都能不用逐个列名指定就完成自动化处理啦~

内容的提问来源于stack exchange,提问作者user9238790

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:15:35