如何按trial分组,基于event_df的索引对提取player_df指定行?
解决方法
步骤1:拆分事件起止索引
先把event_df中存储索引对的indices列拆分为单独的起始索引和结束索引列,方便后续范围筛选:
event_df[['start_idx', 'end_idx']] = pd.DataFrame(event_df['indices'].tolist(), index=event_df.index)
步骤2:编写提取逻辑函数
定义一个函数,针对event_df的每一行(即单个被试的单轮试验),从player_df中筛选出符合条件的行:
def extract_trial_data(row): # 匹配同被试、同试验,且索引在起止范围内的行 filter_condition = (player_df['subject'] == row['subject']) & \ (player_df['trial'] == row['trial']) & \ (player_df.index >= row['start_idx']) & \ (player_df.index <= row['end_idx']) return player_df[filter_condition]
注:如果player_df的索引不是事件对应的目标索引,而是有单独列(比如frame_index),把player_df.index替换为player_df['frame_index']即可。
步骤3:批量提取并合并结果
遍历event_df的每一行执行提取,最后将所有结果合并为一个DataFrame:
final_result = pd.concat(event_df.apply(extract_trial_data, axis=1).tolist(), ignore_index=True)
示例验证
假设我们有以下测试数据:
import pandas as pd # 测试用event_df event_df = pd.DataFrame({ 'subject': ['sub001', 'sub001', 'sub002'], 'trial': [1, 2, 1], 'indices': [[3, 7], [12, 16], [5, 9]] }) # 测试用player_df(行索引对应事件索引) player_df = pd.DataFrame({ 'subject': ['sub001']*20 + ['sub002']*15, 'trial': [1]*10 + [2]*10 + [1]*15, 'x_coord': range(35), 'y_coord': range(35, 70) }) player_df.index.name = 'frame_index'
运行上述步骤后,final_result会包含:
- sub001的trial1中frame_index 3-7的行
- sub001的trial2中frame_index 12-16的行
- sub002的trial1中frame_index 5-9的行
内容的提问来源于stack exchange,提问作者B_Sil
相关产品推荐
相关产品推荐

