R语言:如何基于时间创建符号变化值的直方图?
Let’s break this into two straightforward parts: processing your data to calculate how long each positive/negative state lasts, then visualizing those durations with a histogram. I’ll use Python with pandas and matplotlib—tools that are standard for this kind of data work.
Step 1: Process Your Data to Get State Durations
First, let’s assume your dataset has two columns: timestamp (datetime format) and value (the numerical values you’re tracking). Here’s how to compute the duration of each consecutive sign segment:
- Import the necessary libraries:
import pandas as pd import numpy as np import matplotlib.pyplot as plt
- Load and clean your data:
Make sure your timestamp column is parsed as datetime (if it isn’t already):
df = pd.read_csv("your_data.csv") df['timestamp'] = pd.to_datetime(df['timestamp'])
- Label each value’s sign:
Usenp.sign()to tag values as positive (1), negative (-1), or zero (0). If you want to treat zeros as part of the previous non-zero state, add the optional fill step:
df['sign'] = np.sign(df['value']) # Optional: Fill zeros with the last non-zero sign df['sign'] = df['sign'].replace(0, method='ffill')
- Group consecutive sign segments:
Create a unique ID for each stretch of the same sign—this lets us calculate how long each segment lasts:
# Flag rows where the sign changes from the previous row df['sign_changed'] = df['sign'] != df['sign'].shift(1) # Assign a unique ID to each consecutive sign segment df['segment_id'] = df['sign_changed'].cumsum()
- Calculate duration for each segment:
Group by segment ID and sign, then compute the time difference between the first and last timestamp in each group:
duration_df = df.groupby(['segment_id', 'sign']).agg( start_time=('timestamp', 'first'), end_time=('timestamp', 'last') ).reset_index() # Convert duration to hours (swap to /60 for minutes, or remove division for seconds) duration_df['duration_hours'] = (duration_df['end_time'] - duration_df['start_time']).dt.total_seconds() / 3600
Now you have a dataframe with every sign segment’s duration in hours.
Step 2: Build the Histogram
You can visualize this data in two useful ways—either combine all durations, or split them by sign to compare positive vs negative states.
Option 1: Histogram of All State Durations
plt.figure(figsize=(10, 6)) plt.hist(duration_df['duration_hours'], bins=15, edgecolor='black') plt.title('Distribution of Sign State Durations') plt.xlabel('Duration (Hours)') plt.ylabel('Number of Segments') plt.grid(axis='y', alpha=0.75) plt.show()
Option 2: Side-by-Side Histograms for Positive vs Negative States
positive_durations = duration_df[duration_df['sign'] == 1]['duration_hours'] negative_durations = duration_df[duration_df['sign'] == -1]['duration_hours'] plt.figure(figsize=(12, 6)) plt.hist([positive_durations, negative_durations], bins=15, label=['Positive', 'Negative'], edgecolor='black') plt.title('Positive vs Negative State Duration Distribution') plt.xlabel('Duration (Hours)') plt.ylabel('Number of Segments') plt.legend() plt.grid(axis='y', alpha=0.75) plt.show()
Quick Tips
- Handling Zeros: If zeros should be their own state, skip the
ffillstep and add a label for 0 in your histogram. - Adjust Time Units: Swap the division in the duration calculation to fit your needs (e.g.,
/60for minutes). - Tweak Bins: Play with the
binsparameter inplt.hist()to make the histogram’s detail level match your data.
That’s it—you’ll have a clear histogram showing how long each sign state persisted in your dataset!
内容的提问来源于stack exchange,提问作者Ch Gabri

