如何根据DataFrame首列值将指定列批量设为nan?
按DataFrame首列值动态修改对应列数值
问题说明
需要根据DataFrame第一列col_ind的数值,将每行中从倒数第col_ind列开始到末尾的所有列值设为NaN。
示例输入DataFrame:
import pandas as pd import numpy as np df = pd.DataFrame({ 'col_ind': [3, 2, 1], 'col_1': ['a', 'd', 'g'], 'col_2': ['b', 'e', 'h'], 'col_3': ['c', 'f', 'i'] })
对应表格:
| col_ind | col_1 | col_2 | col_3 |
|---|---|---|---|
| 3 | a | b | c |
| 2 | d | e | f |
| 1 | g | h | i |
期望输出:
| col_ind | col_1 | col_2 | col_3 |
|---|---|---|---|
| 3 | NaN | NaN | NaN |
| 2 | d | NaN | NaN |
| 1 | g | h | NaN |
原代码的问题
你尝试的df.loc[:, df.columns[-df['col_ind']:]] = np.nan无法正常运行,因为df['col_ind']是多行的Series,列索引切片需要单个标量值,不能直接用Series做切片参数。
正确实现方案
方法1:布尔矩阵批量赋值
# 提取需要修改的列(排除首列) target_cols = df.columns[1:] col_count = len(target_cols) # 生成每行的NaN标记矩阵:需要设为NaN的位置标记为True nan_mask = np.array([ np.concatenate([np.zeros(col_count - x, dtype=bool), np.ones(x, dtype=bool)]) for x in df['col_ind'] ]) # 对目标列应用掩码赋值NaN df[target_cols] = df[target_cols].mask(nan_mask)
方法2:逐行处理
def mark_nan(row): nan_col_num = row['col_ind'] # 筛选出当前行需要设为NaN的列(跳过首列) cols_to_nan = df.columns[-nan_col_num:].drop('col_ind', errors='ignore') row[cols_to_nan] = np.nan return row df = df.apply(mark_nan, axis=1)
两种方法都能得到符合预期的输出结果。
内容的提问来源于stack exchange,提问作者thefrollickingnerd
相关产品推荐
相关产品推荐

