You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame选中列间自动插入‘Keep’列?技术求助

在选中列的每两列之间自动插入'Keep'列的解决方案

问题概述

我是编程新手,现在需要处理一个DataFrame,需求是:展示用户选中的列,并且在每两列之间插入内容为'Keep'的列。目前已经实现了列选择功能,但无法自动插入'Keep'列(手动添加除外)。执行代码时报错KeyError: Index(['Keep'], dtype='object'),推测是因为'Keep'不属于原DataFrame的列,其他环节无问题,求解决方案。

原代码与报错信息

列定义

name_of_cols = ['id','start_date', 'end_date', 'name', 'job_title', 'Keep']

除'Keep'外,其余都是原DataFrame的列。

处理函数与调用

def clean_df(df, list_col):
    df2 = df.copy()
    df2 = df2.drop_duplicates(list_col)
    df3 = df2.copy()
    df3 = df3[[id,start_date, end_date, name, job_title]].reset_index(drop = true)
    df_3 = df3_new.columns.tolist()
    conditions =[df3 = name_of_cols,
    df3!= name_of_cols
    results = ['Keep' , 'No Keep']
    df3_new['Keep'] = np.select(conditions, results)
    return df3[name_of_cols]
df3_new = cleanup_df(df3, name_of_cols)

报错信息

KeyError: Index(['Keep'], dtype='object')

错误原因分析

  1. 语法错误:列名未加引号(如id应改为'id')、布尔值true需大写为True、conditions和results的语法结构不完整(缺失括号、逗号)。
  2. 逻辑错误:未定义df3_new就直接调用;np.select的条件写法完全不符合要求,无法生成有效布尔判断。
  3. 核心问题:试图直接返回包含'Keep'的列,但原DataFrame不存在该列,触发KeyError;同时代码逻辑未贴合「在每两列之间插入Keep列」的需求。

解决方案

实现步骤

  1. 从传入的列列表中过滤出原DataFrame存在的列(排除'Keep')。
  2. 构造新列顺序:遍历选中列,每添加一个列后插入'Keep'列(最后一列后不插入)。
  3. 为新增的'Keep'列赋值固定内容'Keep'。
  4. 返回按新列顺序排列的DataFrame。

完整代码示例

import pandas as pd

# 示例原DataFrame(可替换为你的实际数据)
df = pd.DataFrame({
    'id': [1, 2, 3],
    'start_date': ['2023-01-01', '2023-02-01', '2023-03-01'],
    'end_date': ['2023-01-31', '2023-02-28', '2023-03-31'],
    'name': ['Alice', 'Bob', 'Charlie'],
    'job_title': ['Engineer', 'Designer', 'Manager']
})

name_of_cols = ['id','start_date', 'end_date', 'name', 'job_title', 'Keep']

def clean_df(df, list_col):
    # 过滤出原DF中存在的有效列(排除'Keep')
    selected_cols = [col for col in list_col if col != 'Keep']
    # 构建新的列顺序:每两列之间插入'Keep'
    new_col_order = []
    for idx, col in enumerate(selected_cols):
        new_col_order.append(col)
        # 最后一列后不插入Keep
        if idx != len(selected_cols) - 1:
            new_col_order.append('Keep')
    
    df_clean = df.copy().drop_duplicates(selected_cols)
    # 为所有Keep列赋值
    for col in new_col_order:
        if col == 'Keep' and col not in df_clean.columns:
            df_clean[col] = 'Keep'
    
    # 返回按新顺序排列的结果
    return df_clean[new_col_order].reset_index(drop=True)

# 调用函数
df3_new = clean_df(df, name_of_cols)
print(df3_new)

代码说明

  • selected_cols:提取用户选中的有效列,避免无效列干扰。
  • new_col_order:通过循环确保每两个选中列之间插入'Keep',严格贴合需求。
  • 动态创建'Keep'列并赋值,避免因原DataFrame无该列导致的KeyError。

内容的提问来源于stack exchange,提问作者That_non_coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 00:35:35