Pandas技术实操:DataFrame子集访问与峰值异常及局部趋势分析
Alright, let's tackle your Pandas questions step by step with concrete examples—this should make it easy to follow along!
1. Accessing DataFrame Subsets with iterrows(), itertuples(), and Indexing
First, let's create a sample time-series DataFrame to work with:
import pandas as pd import numpy as np # Sample DataFrame with time, X, and Y columns df = pd.DataFrame({ 'time': pd.date_range(start='2023-01-01', periods=30, freq='D'), 'X': np.random.randint(10, 100, size=30), 'Y': np.random.randn(30) })
Index-Based Access (Fastest for Direct Subsets)
This is the most efficient way to grab subsets directly, no iteration needed:
- Label-based with
.loc[]: Use row labels (or boolean masks)# Grab rows where X > 70 high_x_subset = df.loc[df['X'] > 70] # Grab rows between index 5 and 15 (inclusive) range_subset = df.loc[5:15] - Position-based with
.iloc[]: Use integer positions (0-indexed)# Grab first 10 rows first_10 = df.iloc[:10] # Grab rows 5 to 10 (exclusive of 10) and columns 0 and 2 specific_subset = df.iloc[5:10, [0,2]]
iterrows(): Iterate with Row as Series
Great for row-by-row processing where you need to reference columns by name, but note it's slower for large datasets:
for index, row in df.iterrows(): # Access values using column names current_x = row['X'] current_time = row['time'] # Example: Grab subset around rows where X is a local peak if index > 0 and index < len(df)-1: if current_x > df['X'].iloc[index-1] and current_x > df['X'].iloc[index+1]: # Get 2 rows before and after the peak peak_subset = df.loc[index-2:index+2] print(f"Peak at index {index}:\n{peak_subset}\n")
itertuples(): Faster Iteration with Named Tuples
This is the preferred method for large datasets—it returns lightweight named tuples instead of Series, so it's much faster:
for row in df.itertuples(): # Access values using attribute names (or index positions) current_x = row.X current_index = row.Index # Same peak subset example if current_index > 0 and current_index < len(df)-1: if current_x > df['X'].iloc[current_index-1] and current_x > df['X'].iloc[current_index+1]: # .iloc is left-exclusive, so use +3 to include index+2 peak_subset = df.iloc[current_index-2:current_index+3] print(f"Peak at index {current_index}:\n{peak_subset}\n")
2. Identifying X Peak Anomalies & Extracting Surrounding Rows
To identify peak anomalies using X and Y's state, let's define clear criteria (adjust these to your actual data!):
- X is a local maximum (higher than adjacent rows)
- X is an outlier (above mean + 2*standard deviation)
- Y shows abnormal behavior (absolute value > 1, for example)
Here's how to implement this:
# 1. Flag local maxima (window=3 checks current vs previous/next row) df['is_local_max'] = df['X'] == df['X'].rolling(window=3, center=True).max() # 2. Flag X outliers x_mean, x_std = df['X'].mean(), df['X'].std() df['is_x_outlier'] = df['X'] > (x_mean + 2 * x_std) # 3. Flag Y anomalies df['is_y_anomaly'] = abs(df['Y']) > 1 # 4. Combine flags to get peak anomalies df['peak_anomaly'] = df['is_local_max'] & df['is_x_outlier'] & df['is_y_anomaly'] # Get indices of all peak anomalies peak_indices = df[df['peak_anomaly']].index.tolist()
Now extract 5 rows before and after each anomaly (handling edge cases to avoid index errors):
anomaly_subsets = [] for idx in peak_indices: # Ensure we don't go out of bounds start_idx = max(0, idx - 5) end_idx = min(len(df)-1, idx + 5) # Grab the subset subset = df.loc[start_idx:end_idx] anomaly_subsets.append(subset) print(f"Anomaly at index {idx} (time: {df.loc[idx, 'time']}) - surrounding rows:\n{subset}\n")
3. Extracting Local Trend Subsequence & Verifying No Reversal
Let's use the first peak anomaly as the start of our local trend. We'll analyze if the trend stays consistent (no reversal) using two practical methods:
Step 1: Extract the Local Trend Subsequence
if peak_indices: trend_start_idx = peak_indices[0] # Extract from the peak to 20 rows after (adjust length as needed) local_trend_df = df.loc[trend_start_idx:trend_start_idx+20].copy() # Add a "days from start" column for trend analysis local_trend_df['days_since_peak'] = (local_trend_df['time'] - local_trend_df['time'].iloc[0]).dt.days
Step 2: Analyze Trend Reversal
Method 1: Linear Regression Slope
A positive slope means upward trend, near-zero means flat, negative means reversal:
from sklearn.linear_model import LinearRegression # Prepare data for regression X = local_trend_df[['days_since_peak']] y = local_trend_df['X'] model = LinearRegression() model.fit(X, y) slope = model.coef_[0] if slope > 0.1: print(f"Local trend slope: {slope:.2f} → Trend is rising (no reversal)") elif abs(slope) <= 0.1: print(f"Local trend slope: {slope:.2f} → Trend is flat (no reversal)") else: print(f"Local trend slope: {slope:.2f} → Trend has reversed (declining)")
Method 2: Rolling Difference Ratio
Check what percentage of time X is increasing:
# Calculate daily change in X local_trend_df['x_daily_change'] = local_trend_df['X'].diff() # Ratio of days where X increased (ignore first NaN) positive_change_ratio = (local_trend_df['x_daily_change'] > 0).sum() / len(local_trend_df['x_daily_change'].dropna()) if positive_change_ratio > 0.6: print(f"X increased in {positive_change_ratio:.2%} of days → Trend likely not reversed") else: print(f"X increased in {positive_change_ratio:.2%} of days → Trend may have reversed")
内容的提问来源于stack exchange,提问作者Life2Day

