获取含指定字符串的列位置并拆分Pandas DataFrame
解决DataFrame按含'TIME'列拆分的问题
修复InvalidIndexError问题
你触发报错是因为df.columns.get_loc()只支持传入单个列名,不能直接传列名列表。要获取多个目标列的索引,用以下两种方法:
- 方法1:使用
get_indexer批量获取索引# 假设time_cols是你筛选出的含'TIME'的列名列表 time_indices = df.columns.get_indexer(time_cols) - 方法2:循环遍历列名逐个获取索引
time_indices = [df.columns.get_loc(col) for col in time_cols]
注意:建议直接从DataFrame列中筛选有效列名,避免无效值:
time_cols = [col for col in df.columns if 'TIME' in col]
拆分DataFrame为子DataFrame
拿到目标列的索引后,按以下步骤拆分:
- 对索引列表排序(确保列的先后顺序正确)
- 在索引列表末尾添加
df.shape[1](作为最后一个子DataFrame的结束边界) - 遍历索引对,通过切片生成子DataFrame
示例代码:
# 确保索引按列顺序排列 time_indices.sort() # 添加最后一个拆分的结束点 split_points = time_indices + [df.shape[1]] # 生成子DataFrame列表 sub_dfs = [] for i in range(len(split_points)-1): start_idx = split_points[i] end_idx = split_points[i+1] # 切片生成从当前TIME列到下一个TIME列前的子DataFrame sub_df = df.iloc[:, start_idx:end_idx] sub_dfs.append(sub_df)
内容的提问来源于stack exchange,提问作者Brain_overflowed
相关产品推荐
相关产品推荐

