使用Pandas线性插值比较采样时间不同的时间序列
Great question! When dealing with time series data that has inconsistent daily start times and fluctuating sampling rates—especially when you need to compare two sets of identical experimental conditions (like your first 3 vs last 3 days)—aligning the data to a common time grid with linear interpolation is the perfect approach. Let's break this down step by step, tailored to your scenario:
Step 1: Load and Prepare Your Data
First, make sure your timestamp column is parsed as a datetime type and set as the DataFrame index. This is non-negotiable for Pandas' time series tools to work correctly.
import pandas as pd # Load your dataset (adjust file path/column names to match your data) df = pd.read_csv("your_experiment_data.csv") # Convert timestamp column to datetime and set as index df["timestamp"] = pd.to_datetime(df["timestamp"]) df = df.set_index("timestamp").sort_index() # Sort index to guarantee chronological order
Step 2: Create a Standardized Time Grid
Since your daily start times are shifted, we need a uniform time axis to align both datasets. Let's define a grid that covers a full 3-day period, starting at midnight (00:00) and using a fixed frequency matching your typical sampling rate (adjust freq to your needs—e.g., '1min' for 1-minute intervals, '30s' for 30 seconds):
# Get the midnight start of your first experimental day start_ref = df.index.min().normalize() # Define the end of the 3-day reference period (1 second before midnight of the 4th day) end_ref = start_ref + pd.Timedelta(days=3) - pd.Timedelta(seconds=1) # Generate the common time grid common_time_grid = pd.date_range(start=start_ref, end=end_ref, freq="1min")
Step 3: Isolate the Two Datasets to Compare
Extract your first 3 days and last 3 days of experimental data:
# Grab the first 3 days of data first_3_days = df.loc[start_ref : start_ref + pd.Timedelta(days=3)] # Grab the last 3 days (adjust the date range to match your experiment's end date) last_3_days_start = df.index.max().normalize() - pd.Timedelta(days=3) last_3_days = df.loc[last_3_days_start : last_3_days_start + pd.Timedelta(days=3)]
Step 4: Apply Linear Interpolation to Align to the Common Grid
Use Pandas' reindex() method with method='linear' to interpolate both datasets onto our standardized time grid. This will fill in missing time points using linear interpolation between adjacent measured values, naturally accounting for your variable sampling rate.
# Interpolate first 3 days to the common time grid first_3_interpolated = first_3_days.reindex(common_time_grid, method="linear") # Interpolate last 3 days to the same common time grid last_3_interpolated = last_3_days.reindex(common_time_grid, method="linear")
Step 5: Align Dates for Direct Comparison
Right now, the last 3 days have different calendar dates than the first 3. To compare them directly, replace the index of the last 3 days with the index from the first 3 days—this gives both datasets the same "relative time" axis (e.g., 00:00 Day 1, 00:01 Day 1, etc.):
# Align last 3 days to the first 3 days' time index last_3_aligned = last_3_interpolated.set_index(first_3_interpolated.index)
Step 6: Analyze and Compare the Data
Now you're ready to directly compare the two sets of experimental conditions! Here are a few common ways to do this:
- Calculate point-by-point differences:
difference = first_3_interpolated - last_3_aligned - Plot both series on the same axis to visualize trends:
import matplotlib.pyplot as plt plt.figure(figsize=(12, 6)) first_3_interpolated["your_metric_column"].plot(label="First 3 Days") last_3_aligned["your_metric_column"].plot(label="Last 3 Days (Aligned)") plt.legend() plt.title("Comparison of Experimental Conditions") plt.xlabel("Relative Time (3-Day Period)") plt.show() - Compute summary statistics (mean, standard deviation, etc.) for each time interval to quantify differences.
Key Tips to Keep in Mind:
- Edge Cases: If you have missing data at the very start/end of either dataset, linear interpolation won't fill those gaps—use the
limitparameter inreindex()or handle these edge points separately if they're critical to your analysis. - Frequency Tuning: Adjust the
freqparameter indate_range()to match your actual sampling resolution (e.g., '5s' for 5-second intervals if that's your typical rate). - Monotonic Index: Always ensure your time index is sorted (we did this with
sort_index()in Step 1)—Pandas interpolation requires a chronological order to work properly.
内容的提问来源于stack exchange,提问作者bicarlsen

