如何基于后续状态计算机器特定状态的停留时间分布并排除故障状态?
Got it, let's tackle this step by step to get the stay time distributions you need for RNG sampling—while excluding the "system fully failed" states (8 and 16):
Step 1: Calculate Stay Duration for Each State
First, we need to compute how long the machine stays in each CurrentState. Since your DataFrame uses datetime as the index, we can leverage time differences between consecutive entries to get durations:
import pandas as pd import numpy as np import matplotlib.pyplot as plt from scipy.stats import expon, norm # Your sample DataFrame df = pd.DataFrame( {'CurrentState': {'2022-02-01 00:20:00': '128', '2022-02-01 00:21:00': '1024', '2022-02-01 00:22:00': '128', '2022-02-01 00:23:00': '1024', '2022-02-01 00:34:00': '128', '2022-02-01 00:35:00': '1024', '2022-02-01 00:37:00': '128', '2022-02-01 00:42:00': '16', '2022-02-01 00:43:00': '8', '2022-02-01 00:47:00': '128'}, 'Next_State': {'2022-02-01 00:20:00': '1024', '2022-02-01 00:21:00': '128', '2022-02-01 00:22:00': '1024', '2022-02-01 00:23:00': '128', '2022-02-01 00:34:00': '1024', '2022-02-01 00:35:00': '128', '2022-02-01 00:37:00': '16', '2022-02-01 00:42:00': '8', '2022-02-01 00:43:00': '128', '2022-02-01 00:47:00': '16'}, 'State_Change': {'2022-02-01 00:20:00': '128-1024', '2022-02-01 00:21:00': '1024-128', '2022-02-01 00:22:00': '128-1024', '2022-02-01 00:23:00': '1024-128', '2022-02-01 00:34:00': '128-1024', '2022-02-01 00:35:00': '1024-128', '2022-02-01 00:37:00': '128-16', '2022-02-01 00:42:00': '16-8', '2022-02-01 00:43:00': '8-128', '2022-02-01 00:47:00': '128-16'}} ) # Ensure index is datetime type df.index = pd.to_datetime(df.index) # Calculate stay duration: next timestamp minus current timestamp df['StayDuration'] = df.index.to_series().diff().shift(-1) # Convert duration to minutes (adjust to seconds/hours if needed) df['StayDurationMinutes'] = df['StayDuration'].dt.total_seconds() / 60 # Drop last row (no next state to calculate duration for) df = df.dropna(subset=['StayDurationMinutes'])
Step 2: Exclude Fully Failed States
Filter out rows where CurrentState is 8 or 16—these are excluded from distribution calculations:
# Define states to exclude FAILED_STATES = {'8', '16'} # Filter DataFrame to keep only valid states filtered_df = df[~df['CurrentState'].isin(FAILED_STATES)]
Step 3: Aggregate Durations by State
Group the filtered data to collect all stay durations for each valid state:
# Group by state and collect duration lists state_duration_groups = filtered_df.groupby('CurrentState')['StayDurationMinutes'].apply(list).to_dict() # Example: Check durations for state 128 print("Stay durations for state 128:", state_duration_groups['128']) # Output: Stay durations for state 128: [1.0, 1.0, 1.0, 5.0, 4.0]
Step 4: Generate Plot-Ready Distribution Objects
We can fit common distributions (exponential, normal) to each state's data and create visualizations to validate fits:
def plot_state_distribution(state, durations): plt.figure(figsize=(8, 5)) # Plot raw data histogram plt.hist(durations, bins='auto', density=True, alpha=0.6, color='g', label='Raw Stay Times') # Fit and plot exponential distribution (common for machine state times) exp_params = expon.fit(durations) x = np.linspace(min(durations), max(durations), 100) plt.plot(x, expon.pdf(x, *exp_params), 'r-', lw=2, label=f'Exponential Fit\nλ={exp_params[1]:.2f}') # Fit and plot normal distribution norm_params = norm.fit(durations) plt.plot(x, norm.pdf(x, *norm_params), 'b-', lw=2, label=f'Normal Fit\nμ={norm_params[0]:.2f}, σ={norm_params[1]:.2f}') plt.title(f'Stay Time Distribution for State {state}') plt.xlabel('Duration (Minutes)') plt.ylabel('Density') plt.legend() plt.show() # Plot distribution for each valid state for state, durations in state_duration_groups.items(): plot_state_distribution(state, durations)
Step 5: Use Distributions for RNG Sampling
Once you have fitted parameters, you can sample from the distributions to simulate stay times in your machine run simulation:
# Example: Sample 10 stay times from state 128's exponential distribution state_128_exp_params = expon.fit(state_duration_groups['128']) sampled_times = expon.rvs(*state_128_exp_params, size=10) print("Sampled stay times for state 128 (exponential):", sampled_times.round(2))
Key Notes:
- Edge Cases: The last row is dropped because there's no subsequent state to calculate a duration. If you need to handle final states, define a custom termination rule for their duration.
- Distribution Choice: Machine state stay times often follow exponential distributions (due to the memoryless property), but validate fits with statistical tests (like Kolmogorov-Smirnov) if needed.
- Units: Adjust duration units (seconds, hours) to match your simulation's requirements.
内容的提问来源于stack exchange,提问作者GT1992

