pandas使用to_string时如何为不同列设置不同的NaN替换值
pandas 1.2.0+ 不同列差异化处理NaN后输出字符串的解决方案
从pandas 1.2.0版本修复NaN不传入格式化函数的问题后,可通过以下两种常用方案实现需求:
方案1:预先逐列替换NaN
该方案逻辑简单直观,适合绝大多数场景,操作时可先基于原数据生成副本避免修改原始数据:
import pandas as pd import numpy as np # 示例数据 df = pd.DataFrame({ 'col1': [1, np.nan, 3, np.nan], 'col2': ['x', np.nan, 'y', np.nan], 'col3': [22.5, np.nan, 45.8, np.nan] }) # 生成输出用副本,逐列替换NaN df_output = df.copy() df_output['col1'] = df_output['col1'].fillna('-') df_output['col2'] = df_output['col2'].fillna('?') df_output['col3'] = df_output['col3'].fillna('') # 输出字符串 print(df_output.to_string(index=False))
方案2:自定义formatters参数处理
如果不想额外生成数据副本,可在to_string的formatters参数中为每列自定义格式化函数,手动处理NaN判断逻辑:
# 定义每列格式化规则,优先判断是否为NaN formatters = { 'col1': lambda x: '-' if pd.isna(x) else f'{x:.0f}', 'col2': lambda x: '?' if pd.isna(x) else x, 'col3': lambda x: '' if pd.isna(x) else f'{x:.1f}' } # 注意必须将na_rep设置为空字符串,避免pandas全局替换NaN覆盖自定义逻辑 print(df.to_string(formatters=formatters, na_rep='', index=False))
如果需要处理的列数较多,可批量生成formatters字典,无需逐个手写规则:
# 提前配置每列对应的NaN替换值 na_rule = { 'col1': '-', 'col2': '?', 'col3': '' } formatters = {} for col, rep_val in na_rule.items(): if df[col].dtype.kind in ('i', 'f'): # 数值类列自定义数值格式化规则 formatters[col] = lambda x, rep=rep_val: rep if pd.isna(x) else f'{x:.1f}' else: # 非数值类列直接返回原值 formatters[col] = lambda x, rep=rep_val: rep if pd.isna(x) else str(x)
内容的提问来源于stack exchange,提问作者Michel de Ruiter
相关产品推荐
相关产品推荐

