如何用Python Pandas基于前次试验条件排序数据及统计特定试验次数
Got it, let's walk through how to tackle these stats step by step using pandas. I’m assuming your DataFrame has columns like choice (storing selections 1/2) and result (I’ll use True for correct results here—swap in strings like "correct" if that's what your data uses).
1. 统计「选择为1且结果正确」的总试验次数
You mentioned you already have this covered, but here’s a standard, concise way to calculate it (for reference and to align with the next steps):
# 两种等价写法,选你习惯的就行 total_correct_1 = ((df['choice'] == 1) & (df['result'] == True)).sum() # 或者用len()筛选后的子集 total_correct_1 = len(df[(df['choice'] == 1) & (df['result'] == True)])
2. 分析「前一次试验满足选1且正确」的后续试验
To pull these stats, we first need to flag which rows are follow-ups to a successful "choice=1" trial. We’ll use pandas' shift() method to reference the previous row’s values.
Step 1: Create a flag column for target follow-up trials
Add a boolean column that marks whether the current row comes right after a "choice=1 and result=correct" trial:
# Shift the condition one row down to mark follow-up rows df['is_follow_up_to_correct_1'] = (df['choice'].shift(1) == 1) & (df['result'].shift(1) == True)
The first row will have NaN here (since there’s no prior trial), and pandas will automatically ignore it in our subsequent calculations—no extra cleanup needed.
Step 2: Calculate your two target counts
Now we can filter using our flag column to get the numbers you need:
# i) 后续试验中「选择为1且结果正确」的次数 follow_up_correct_1 = ((df['is_follow_up_to_correct_1']) & (df['choice'] == 1) & (df['result'] == True)).sum() # ii) 后续试验中「选择为2且结果错误」的次数 follow_up_wrong_2 = ((df['is_follow_up_to_correct_1']) & (df['choice'] == 2) & (df['result'] == False)).sum()
If you prefer using len() for readability, you can rewrite these as:
follow_up_correct_1 = len(df[df['is_follow_up_to_correct_1'] & (df['choice'] == 1) & (df['result'] == True)]) follow_up_wrong_2 = len(df[df['is_follow_up_to_correct_1'] & (df['choice'] == 2) & (df['result'] == False)])
Quick Notes
- If your
resultcolumn uses string values (like "correct"/"incorrect"), just replaceTrue/Falsewith the matching strings in all conditions. - If you don’t want to keep the flag column in your DataFrame long-term, you can chain the shift directly into your filters instead of creating a new column—though the flag makes the logic easier to debug.
内容的提问来源于stack exchange,提问作者jig

