You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按行拆分Excel文件?现有拆分代码遗漏剩余行求修复

问题排查与修复

核心问题分析

  1. 剩余数据覆盖已有文件:循环结束后i的值为13(range(n_chunks)生成0到n_chunks-1的索引),此时保存剩余数据时复用i作为文件名索引,会直接覆盖第13个文件,而非生成新的第14个文件。
  2. ID生成硬编码不健壮:前14个文件的ID用range(1,40)硬编码,若后续修改rows_per_file参数会直接出错,需统一按子DataFrame长度生成。
  3. 路径转义与变量名错误:单反斜杠路径会触发Python转义逻辑;os.walk的第三个参数应为复数filenames,原代码单数写法导致无法正确遍历文件。

修复后的完整代码

import os
import glob
import pandas as pd

rows_per_file = 39
total_rows = len(df)
n_chunks = total_rows // rows_per_file

# 处理完整数据块
for i in range(n_chunks):
    start = i * rows_per_file
    stop = (i + 1) * rows_per_file
    sub_df = df.iloc[start:stop]
    sub_df['ID'] = range(1, len(sub_df) + 1)
    # 用原始字符串避免路径转义问题
    sub_df.to_excel(r"RESULT\{}.xlsx".format(i), sheet_name="Hoja1", index=False, header=True)

# 处理剩余数据块
remaining_rows = total_rows % rows_per_file
if remaining_rows > 0:
    start = n_chunks * rows_per_file
    sub_df = df.iloc[start:]
    sub_df['ID'] = range(1, len(sub_df) + 1)
    # 剩余块使用n_chunks作为索引(即14),避免覆盖已有文件
    sub_df.to_excel(r"RESULT\{}.xlsx".format(n_chunks), sheet_name="Hoja1", index=False, header=True)

# 修正后的文件遍历逻辑
path = r"RESULT"
ext = "*.xlsx"
files = []

for dirpath, dirnames, filenames in os.walk(path):
    files += glob.glob(os.path.join(dirpath, ext))

修复效果

现在会生成14个39行的文件(索引0~13)和1个25行的文件(索引14),总计15个文件,完全符合预期拆分逻辑。

内容的提问来源于stack exchange,提问作者Oliver Briceño

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 11:12:03