求助:基于振荡值为DataFrame标记状态列的实现方法
Alright, let's break down how to solve this problem step by step. From your description, we need to mark alternating positive/negative cycles in a DataFrame, where each cycle starts when column Y transitions from 0 to a positive value, ends at the next 0 (right before it would switch to negative), and repeats this pattern for negative cycles afterward.
Step 1: Prepare Sample Data
First, let's create a sample DataFrame to demonstrate the logic. This mimics the oscillating pattern you described:
import pandas as pd import numpy as np # Set seed for reproducibility np.random.seed(42) # Generate X values and oscillating Y values (0 → positive → 0 → negative → 0...) x = np.linspace(0, 20, 100) y = np.where( (x >= 0) & (x < 5), np.random.uniform(0.5, 2, 25), np.where( (x >= 5) & (x < 10), 0, np.where( (x >= 10) & (x < 15), np.random.uniform(-2, -0.5, 25), 0 ) ) ) df = pd.DataFrame({'X': x, 'Y': y})
Step 2: Identify Cycle Boundaries
We need to detect when Y transitions between 0 and non-zero values, which marks the start/end of each cycle. Here's how to do it:
# 1. Label Y's current state (positive, negative, zero) df['y_state'] = np.where(df['Y'] > 0, 'pos', np.where(df['Y'] < 0, 'neg', 'zero')) # 2. Find indices where the state changes (e.g., zero → pos, pos → zero) state_changes = df['y_state'] != df['y_state'].shift() change_points = df[state_changes].index.tolist()
Step 3: Define and Label Each Cycle
Now we'll iterate through the state change points to define each cycle's start, end, and type (positive/negative), then apply these labels to the DataFrame:
# Initialize list to store cycle details and label column cycles = [] df['cycle_label'] = 'none' # Iterate through state changes to map cycles for i in range(len(change_points) - 1): start_idx = change_points[i] end_idx = change_points[i+1] start_state = df.loc[start_idx, 'y_state'] end_state = df.loc[end_idx, 'y_state'] # Check if this is a valid positive cycle (zero → pos → zero) if start_state == 'zero' and df.loc[start_idx+1, 'y_state'] == 'pos' and end_state == 'zero': cycle_name = f"positive_cycle_{len(cycles)+1}" df.loc[start_idx:end_idx, 'cycle_label'] = cycle_name cycles.append({'type': 'positive', 'start': start_idx, 'end': end_idx}) # Check if this is a valid negative cycle (zero → neg → zero) elif start_state == 'zero' and df.loc[start_idx+1, 'y_state'] == 'neg' and end_state == 'zero': cycle_name = f"negative_cycle_{len(cycles)+1}" df.loc[start_idx:end_idx, 'cycle_label'] = cycle_name cycles.append({'type': 'negative', 'start': start_idx, 'end': end_idx})
Step 4: Verify the Result
You can check the labeled DataFrame with:
# Print first 30 rows (covers first positive cycle and zero transition) print(df[['X', 'Y', 'cycle_label']].head(30)) # Print last 30 rows (covers negative cycle and final zero) print(df[['X', 'Y', 'cycle_label']].iloc[70:100])
Edge Case Handling
- If your dataset starts with a non-zero Y value (not a zero → positive transition), the first partial cycle will be labeled as 'none' (you can adjust this if needed to mark it as incomplete).
- If the dataset ends with a non-zero Y value (no closing zero), that cycle will remain unlabeled; you can extend the logic to mark it as
incomplete_{type}_cycleif required.
内容的提问来源于stack exchange,提问作者Life2Day

