You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何消除Pandas中DataFrame.to_string()列间多余空格并保留固定行宽

解决Pandas to_string()导致固定长度文件行宽异常的问题

我有一个每行固定93字符长度的adb.dat文件,使用Pandas的DataFrame.to_string()方法将数据写入新文件后,每行长度变成了94——该方法会为列对齐自动添加多余空格。先后尝试两种方案都未达到预期:

  • 参考pandas-dev Issue#571设置justify='right',但列间仍存在多余空格,行宽依旧为94;
  • 使用df.iloc[i].fillna('').str.strip().str.cat(sep='')消除空格,虽然移除了多余空格,但行宽缩短至61,不符合固定长度要求。

期望实现:消除列间多余空格,同时保持每行93字符的固定长度,完全还原原文件格式。


相关代码与尝试结果

初始文件adb.dat内容:

20230731;                                                                                     
7914785201501236AAA    TEST TEST                     94481873                                
7914785201502341AAA    TEST TEST                     94481873                                
7914785201503455AAA    TEST TEST                     94481873                                
6736705201501232AAA    TEST TEST                     94481873                                
6736705201502347AAA    TEST TEST                     94481873                                
00000005; 

第一种尝试代码:

df= pd.read_fwf(py_file_fullnm,
skiprows=1,
skipfooter=1,
header=None,
names=['content1','content2'],
colspecs=[(0,16),(16,93)],
delimiter="\0",
skipinitialspace =True,
dtype='string')  

data=df.to_string(index=False,header=None,justify='right')
print(data)
with open('testing.dat','w') as editor:
    editor.write("20230731;  ")
    editor.write("\n"+data)
    editor.write("\n00000007;   ")

第一种尝试结果:

20230731;                                                                                     
7914785201501236 AAA    TEST TEST                     94481873                                
7914785201502341 AAA    TEST TEST                     94481873                                
7914785201503455 AAA    TEST TEST                     94481873                                
6736705201501232 AAA    TEST TEST                     94481873                                
6736705201502347 AAA    TEST TEST                     94481873                                
00000005; 

第二种尝试代码:

df= pd.read_fwf(py_file_fullnm,
skiprows=1,
skipfooter=1,
header=None,
names=['content1','content2'],
colspecs=[(0,16),(16,93)],
delimiter="\0",
skipinitialspace =True,
dtype='string')

with open('testing.dat','w') as editor:
    editor.write("20230731;  ")
    for i in range (0,5):
        data=df.iloc[i].fillna('').str.strip().str.cat(sep='')
        editor.write("\n"+data)
    editor.write("\n00000005;   ")

第二种尝试结果:

20230731;  
7914785201501236AAA    TEST TEST                     94481873
7914785201502341AAA    TEST TEST                     94481873
7914785201503455AAA    TEST TEST                     94481873
6736705201501232AAA    TEST TEST                     94481873
6736705201502347AAA    TEST TEST                     94481873
00000005;   

期望结果:

20230731;                                                                                     
7914785201501236AAA    TEST TEST                     94481873                                
7914785201502341AAA    TEST TEST                     94481873                                
7914785201503455AAA    TEST TEST                     94481873                                
6736705201501232AAA    TEST TEST                     94481873                                
6736705201502347AAA    TEST TEST                     94481873                                
00000005; 

解决方案

核心思路是:既然原文件是固定宽度格式,读取时严格按列位置拆分,输出时直接拼接列内容,避免使用to_string()自动添加列间空格,同时确保每行长度补全到93字符。

import pandas as pd

py_file_fullnm = 'adb.dat'

# 读取固定宽度文件,严格保留原列的所有内容(包括空格)
df = pd.read_fwf(
    py_file_fullnm,
    skiprows=1,
    skipfooter=1,
    header=None,
    names=['content1', 'content2'],
    colspecs=[(0, 16), (16, 93)],  # 与原文件列范围完全一致
    dtype='string'
)

with open('testing.dat', 'w') as editor:
    # 写入开头行,注意保留原行的所有空格
    editor.write("20230731;                                                                                     \n")
    # 遍历每一行,直接拼接列内容,确保长度为93
    for _, row in df.iterrows():
        # 直接拼接两列,不做strip操作,保留原列内的空格
        line = row['content1'] + row['content2']
        # 确保行长度严格为93,不足补空格(读取时已按93截取,此处为保险)
        line = line.ljust(93)
        editor.write(line + "\n")
    # 写入结尾行
    editor.write("00000005; ")

方案说明:

  1. 读取阶段:colspecs严格对应原文件的列位置,确保每列的内容(包括内部的空格)都完整读取,不会丢失或截断;
  2. 输出阶段:直接拼接两列内容,跳过to_string()的自动格式化逻辑,避免添加多余的列间空格;
  3. 长度保证:使用ljust(93)确保每行最终长度为93字符,完全匹配原文件的固定格式。

内容的提问来源于stack exchange,提问作者John Ng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 22:35:09