You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于列值变化分割Python Pandas DataFrame的技术问询

Split Pandas DataFrame into Sub-DataFrames When Column Values Change

Got it, let's walk through how to split your DataFrame into smaller sub-DataFrames whenever a specific column's value shifts. I'll use your sample structure to make this concrete.

Step 1: Create a Grouping Key

First, we need a way to mark rows that belong to the same "continuous block" of values. For example, if we want to split based on the Sign column, we can generate a unique ID for each consecutive group of identical values:

import pandas as pd

# Recreate your sample DataFrame (with a small tweak to show a value change)
data = {
    'Image': [
        'IMG_170705_121224_0148_GRE_vig_ortho_correct.tif',
        'IMG_170705_121226_0149_GRE_vig_ortho_correct.tif',
        'IMG_170705_121228_0150_GRE_vig_ortho_correct.tif',
        'IMG_170705_121230_0151_GRE_vig_ortho_correct.tif',
        'IMG_170705_121232_0152_GRE_vig_ortho_correct.tif',
        'IMG_170705_121234_0153_GRE_vig_ortho_correct.tif',
        'IMG_170705_121236_0154_GRE_vig_ortho_correct.tif',
        'IMG_170705_121238_0155_GRE_vig_ortho_correct.tif'
    ],
    'Yaw': [-41.299461, -39.885353, -38.424816, -44.121506, -43.348404, -42.511200, 10.123456, 11.789012],
    'Sign': [-1.0, -1.0, -1.0, -1.0, -1.0, -1.0, 1.0, 1.0]
}

df = pd.DataFrame(data)

# Generate a group ID: increments every time 'Sign' differs from the previous row
df['group_id'] = df['Sign'].ne(df['Sign'].shift()).cumsum()

Step 2: Split into Sub-DataFrames

Now use groupby to split the original DataFrame into chunks based on our group_id column. You can store these sub-DataFrames in a list or dictionary for easy access:

# Option 1: Store sub-DataFrames in a list
sub_dfs = [group for _, group in df.groupby('group_id')]

# Access the first sub-DataFrame (all rows where Sign = -1.0)
print(sub_dfs[0])

# Option 2: Store in a dictionary (keys are group IDs, values are sub-DataFrames)
sub_dfs_dict = {gid: group for gid, group in df.groupby('group_id')}

# Access the second sub-DataFrame (all rows where Sign = 1.0)
print(sub_dfs_dict[2])

Customize for Other Conditions

If you want to split based on a different rule (like a threshold change in Yaw), adjust how you generate the group_id:

# Split when the absolute change in Yaw is greater than 5
df['yaw_diff'] = df['Yaw'].diff().abs()
df['group_id'] = (df['yaw_diff'] > 5).cumsum()

# Split into sub-DataFrames
yaw_based_sub_dfs = [group for _, group in df.groupby('group_id')]

This approach works for any column or condition—just tweak the group_id logic to match what counts as a "change" for your use case.

内容的提问来源于stack exchange,提问作者Landini135

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:35:27