Pandas处理交替两值列:定位特定连续值索引及常见操作咨询
Got it, let's break down how to solve your specific problem first, then cover common techniques for handling this kind of sequential data in Pandas.
Solution to Your Specific Requirement
First, let's replicate your example data and walk through the steps:
Step 1: Prepare the Data
import pandas as pd import numpy as np # Your example sequence (I=Increase, D=Decrease) seq = ["I","I","I","I","D","I","I","D","I","D","I","D","D","D","D","I","D","I","D","D","I","I","I","I"] df = pd.DataFrame({"change": seq}) # Optional: Set your quarter index if needed # df.index = pd.date_range(start='2000Q1', periods=len(seq), freq='Q')
Step 2: Identify Consecutive Value Groups
We first group consecutive identical values and calculate their lengths:
# Mark groups of consecutive identical values df['group_id'] = (df['change'] != df['change'].shift()).cumsum() # Get summary info for each group: value, start/end index, length group_summary = df.groupby('group_id').agg( value=('change', 'first'), start_idx=('change', 'first_valid_index'), end_idx=('change', 'last_valid_index'), length=('change', 'size') ).reset_index(drop=True)
Step 3: Locate the Target Index
We need to find the first group of at least 2 consecutive Increases that comes after at least one group of 2+ consecutive Decreases:
target_group = None has_valid_decrease_group = False # Iterate through valid groups (length ≥2) to find our target for _, row in group_summary[group_summary['length'] >=2].iterrows(): if row['value'] == 'D': has_valid_decrease_group = True elif row['value'] == 'I' and has_valid_decrease_group: target_group = row break # Get the index of the second Increase in this group (matches your example's result of 21) if target_group is not None: result_index = target_group['start_idx'] + 1 print(f"Target index: {result_index}") # If you have a quarter index, you can get it like this: # print(f"Corresponding quarter: {df.index[result_index]}")
Running this on your example will return 21, which matches your expected result.
Common Techniques for Handling Sequential Value Data in Pandas
Here are some go-to methods for similar tasks you might need later:
1. Count Consecutive Occurrences
Calculate how many times the current value has appeared consecutively up to each row:
# Method 1: Using groupby + cumcount df['consec_count'] = df.groupby('group_id').cumcount() + 1 # Method 2: Using np.where for a more concise approach df['consec_count'] = np.where(df['change'] == df['change'].shift(), df['consec_count'].shift() + 1, 1)
2. Detect Value Change Points
Find where the value switches from Increase to Decrease (or vice versa):
# Mark rows where the value changes from the previous row df['is_change'] = df['change'] != df['change'].shift() # Get the indices of all change points change_indices = df[df['is_change']].index.tolist()
3. Detect N Consecutive Occurrences of a Value
Check if a value appears N times in a row (e.g., 3 consecutive Decreases):
n = 3 # Using rolling window df[f'consec_{n}_d'] = df['change'].rolling(n).apply(lambda x: all(x == 'D')).astype(bool) # Using shift (faster for small N) df[f'consec_{n}_d'] = (df['change'] == 'D') & (df['change'].shift() == 'D') & (df['change'].shift(2) == 'D')
4. Find First Occurrence After a Specific Pattern
For example, find the first Increase immediately after 2 consecutive Decreases:
# Mark rows where 2 consecutive Decreases end df['two_d_end'] = (df['change'] == 'D') & (df['change'].shift() == 'D') # Get the first Increase after the first valid two_d_end first_i_after_two_d = df.loc[df['two_d_end'].idxmax():, 'change'].eq('I').idxmax()
内容的提问来源于stack exchange,提问作者CuriousLearner

