You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用更优雅方式替换Python中清洗DataFrame列的三层嵌套if-else

优化Pandas DataFrame列清洗代码

需求说明

需要实现以下DataFrame列清洗逻辑:

  • 补充缺失列并填充空字符串
  • 将全NaN列转为空字符串
  • 所有列转为字符串格式,NaN替换为空字符串
  • 指定列(如total)保留浮点格式字符串,其他列转为整数格式字符串

原实现使用三层嵌套if-else,逻辑繁琐,以下是更简洁优雅的优化方案:


优化后完整代码

import numpy as np
import pandas as pd

# 示例数据
data = [[1, 2.4, 3, np.nan], [4, 5.3, 6, np.nan], [np.nan, 8, 3, np.nan]]
output_data = pd.DataFrame(data, columns=['total', 'count1', 'count2', 'count3'])

# 配置参数
all_cols = ['total', 'count1', 'count2', 'count3', 'count4']
float_cols = {'total'}  # 用集合提升成员查找效率

# 1. 补充缺失列并填充空字符串
missing_cols = set(all_cols) - set(output_data.columns)
output_data[list(missing_cols)] = ''

# 2. 定义列格式化函数,拆分嵌套逻辑
def format_col(col_series, is_float_col):
    if col_series.isna().all():
        return col_series.fillna('')
    filled_series = col_series.fillna('')
    if is_float_col:
        return filled_series.astype(str)
    # 非浮点列截取整数部分
    return filled_series.astype(str).str.split('.').str[0]

# 3. 批量处理目标列
for col in all_cols:
    if col not in missing_cols:
        output_data[col] = format_col(output_data[col], col in float_cols)

# 查看结果
print(output_data)

优化亮点

  • 集合操作简化缺失列处理:用集合差集快速定位缺失列,替代循环判断,代码更简洁且效率更高
  • 函数拆分嵌套逻辑:将原三层嵌套的判断逻辑提取为独立函数,可读性大幅提升,后续修改规则也更方便
  • 集合提升查找效率:将float_cols从列表改为集合,成员查找时间复杂度从O(n)降至O(1)
  • 冗余逻辑复用:将重复的空值填充逻辑统一到函数中,减少代码重复

输出结果

total count1 count2 count3 count4
0   1.0      2      3             
1   4.0      5      6             
2            8      3             

内容的提问来源于stack exchange,提问作者Canovice

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 04:32:38