You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas读取带末尾空行文本时DataFrame列表多余逗号解决方法

pandas读取带末尾空行文本生成列表列出现多余逗号的修复方案

问题复现

使用pandas批量读取目录下文本文件时,若文件末尾存在空换行,将换行替换为逗号再分割生成列表列后,列表尾部会出现多余空值,表现为多余逗号。
复现代码如下:

from pathlib import Path
import pandas as pd

files = Path('data/').glob('*')

df = list()

for file in files:
    df.append(file.read_text().replace('\n', ','))  # 直接替换所有换行符为逗号

df = pd.DataFrame(df, columns = ['fruits'])
df = df['fruits'].str.split(',').to_frame()
df

测试用文件:

  • 文件1:包含apple/banana/orange3行内容,末尾带2个空换行
  • 文件2:包含kiwi/mango/grapes/berry/coconut5行内容,末尾无空换行
    异常运行输出:
fruits
0  [kiwi, mango, grapes, berry, coconut]
1           [apple, banana, orange, ,]

期望输出:

fruits
0  [kiwi, mango, grapes, berry, coconut]
1               [apple, banana, orange]

问题原因

read_text()方法会完整读取文件内所有字符,包括末尾的连续空换行符。直接将所有\n替换为逗号时,末尾的换行符会被转为多余的尾部逗号,后续按逗号分割时就会生成空字符串元素,最终表现为列表尾部的多余逗号。

修复方案

方案1:读取时移除首尾空换行(性能最优)

读取文本内容后,先用strip('\n')移除字符串首尾所有连续的换行符,再执行替换逻辑,仅需修改循环内1行代码,无额外计算开销。
注意:strip('\n')仅处理字符串首尾的连续换行,不会改动正常行内容之间的换行,对末尾无空行的文件无任何影响
修正后完整代码:

from pathlib import Path
import pandas as pd

files = Path('data/').glob('*')

df = list()

for file in files:
    # 新增strip('\n')提前移除首尾空换行,避免生成多余逗号
    df.append(file.read_text().strip('\n').replace('\n', ','))

df = pd.DataFrame(df, columns = ['fruits'])
df = df['fruits'].str.split(',').to_frame()
df

方案2:分割后过滤空元素(鲁棒性更强)

如果需要兼容文件中存在的其他意外空行(如开头空行、内容中间夹杂空行),可以在分割步骤后过滤掉所有空字符串元素,适配更多异常文件场景:

from pathlib import Path
import pandas as pd

files = Path('data/').glob('*')

df = list()

for file in files:
    df.append(file.read_text().replace('\n', ','))

df = pd.DataFrame(df, columns = ['fruits'])
# 分割后过滤所有空字符串元素
df = df['fruits'].str.split(',').apply(lambda x: [i for i in x if i]).to_frame()
df

两种方案运行后均可得到符合预期的输出结果。

内容的提问来源于stack exchange,提问作者nerd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 22:51:19