You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas新手求助:按全0行拆分DataFrame(需for循环实现)

按全0行拆分Pandas DataFrame(for循环实现)

步骤说明

  1. 识别所有整行值均为0的行
  2. 生成拆分的边界索引
  3. 通过for循环遍历边界,提取非空的子DataFrame

完整代码示例

import pandas as pd

# ----------------------
# 1. 构造模拟数据(替换成你的真实DataFrame)
# ----------------------
data = {
    'A': [1, 2, 0, 0, 3, 4, 0, 5],
    'B': [2, 3, 0, 0, 4, 5, 0, 6],
    'C': [3, 4, 0, 0, 5, 6, 0, 7]
}
df = pd.DataFrame(data)

# ----------------------
# 2. 核心拆分逻辑
# ----------------------
# 标记每行是否全为0
is_zero_row = df.eq(0).all(axis=1)
# 获取全0行的索引,转为列表
zero_indices = df.index[is_zero_row].tolist()
# 补充拆分的首尾边界:从第0行开始,到最后一行结束
split_bounds = [0] + zero_indices + [len(df)]

# 用列表存储拆分后的子DataFrame(推荐,方便批量操作)
split_dfs = []
for i in range(len(split_bounds) - 1):
    start_idx = split_bounds[i]
    end_idx = split_bounds[i+1]
    # 提取子DataFrame
    sub_df = df.iloc[start_idx:end_idx]
    # 跳过空的子DataFrame,以及全是0的子DataFrame
    if not sub_df.empty and not sub_df.eq(0).all(axis=1).all():
        split_dfs.append(sub_df)

# ----------------------
# 3. (可选)将子DataFrame命名为df1、df2...(直接赋值到全局变量)
# ----------------------
for idx, sub_df in enumerate(split_dfs, start=1):
    globals()[f'df{idx}'] = sub_df

# 测试输出
print("拆分后的df1:")
print(df1)
print("\n拆分后的df2:")
print(df2)
print("\n拆分后的df3:")
print(df3)

关键细节解释

  • df.eq(0).all(axis=1):逐行检查所有列的值是否都等于0,返回布尔Series,True代表该行全为0
  • split_bounds:把全0行的索引作为拆分点,再加上起始的0和结束的总行数,以此确定每个子DataFrame的起止区间
  • 循环中跳过空DataFrame:避免连续全0行导致生成空的子DataFrame,同时排除本身就是全0的子DataFrame

注意事项

  • 如果你的DataFrame使用了非默认整数索引,建议先执行df = df.reset_index(drop=True)重置索引,避免切片出错
  • 若数据量极大,优先用split_dfs列表存储子DataFrame,比动态生成df1、df2这类变量更高效、易维护
  • 该方法基于Pandas的索引切片,处理大行数DataFrame时性能稳定

内容的提问来源于stack exchange,提问作者Zlatan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 15:57:50