如何在Python中按特定行的第四列值拆分DataFrame生成子数据帧
问题描述
我读取文本文件生成了如下结构的DataFrame:
import pandas as pd filename = r"/home/EVENTS.S001.R001" # 返回仅含一列的DataFrame df = pd.read_table(filename, header=None) df = df[0].str.split(" ", expand=True)
对应的DataFrame输出示例:
0 2.0 0 5 6.31000 None None None None None None None None None None None 1 5.60196 2.98423 2.21817 0.04454 0.00000 None None None None None None 2 2.0 0 5 6.31000 None None None None None None None None None None None 3 5.23222 0.87787 0.67289 0.04454 0.00000 None None None None None None 4 2.0 0 4 6.31000 None None None None None None None None None None None ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... 200024 3.92921 1.01950 0.04454 0.00000 None None None None None None None None 200025 2.0 0 3 6.31000 None None None None None None None None None None None 200026 3.70093 1.50100 0.00000 None None None None None None None None None None 200027 2.0 0 3 6.31000 None None None None None None None None None None None 200028 4.35257 1.63800 0.00000 None None None None None None None None None None
需求:按所有以2.0开头的行的第四列值(如第一行该值为5)分类,将每个值对应的下一行数据提取出来,生成对应子DataFrame。例如值为5时,子DataFrame如下:
5.60196 2.98423 2.21817 0.04454 5.23222 0.87787 0.67289 0.04454 0.00000
需要为该列取值1到10分别生成子DataFrame。
解决方案
通过以下步骤实现需求:
- 标记目标行:筛选出所有以
2.0开头的行,记录索引并提取第四列的分类值(转为整数,匹配1-10的范围)。 - 提取对应下一行数据:根据标记的索引,取每个索引+1的行作为对应分类的数据行。
- 整理并分类存储:清理数据行中的
None值,按分类值分组生成子DataFrame,用字典存储方便调用。
完整代码如下:
import pandas as pd filename = r"/home/EVENTS.S001.R001" df = pd.read_table(filename, header=None) df = df[0].str.split(" ", expand=True) # 1. 筛选以2.0开头的行,提取分类值并限定1-10范围 target_rows = df[df[0] == "2.0"].copy() target_rows["category"] = target_rows[3].astype(int) target_rows = target_rows[target_rows["category"].between(1, 10)] # 2. 遍历目标行,提取对应下一行数据并分类存储 result_dict = {} for idx, row in target_rows.iterrows(): cat = row["category"] # 获取下一行,清理None值并重置列索引 data_row = df.loc[idx+1].dropna().reset_index(drop=True) # 将数据添加到对应分类的列表中 if cat not in result_dict: result_dict[cat] = [] result_dict[cat].append(data_row) # 3. 将每个分类的列表转换为DataFrame for cat in result_dict: result_dict[cat] = pd.DataFrame(result_dict[cat]) result_dict[cat].reset_index(drop=True, inplace=True) # 调用示例:获取分类为5的子DataFrame print(result_dict[5])
代码说明
- 第一步筛选有效目标行,仅保留分类值在1-10范围内的数据,避免无效处理。
- 提取下一行数据时,用
dropna()移除None值,确保子DataFrame仅包含有效数据。 - 最终用字典
result_dict存储所有子DataFrame,键为分类值(1-10),值为对应子DataFrame,可直接通过键调用。
内容的提问来源于stack exchange,提问作者Montana
相关产品推荐
相关产品推荐

