优化Pandas时间索引DataFrame新增滞后列的低效函数
Hey there! Let's fix that slow performance issue you're seeing with your nn_format_df function. The core problem with most slow Pandas operations is that they rely on row-by-row iteration (like iterrows() or custom loops), which completely bypasses Pandas' optimized vectorized operations. Here's a clean, Pandas-native solution that'll get your runtime down to the sub-1-second range:
The Fast, Vectorized Approach
Since your DataFrame uses a datetime index, we can leverage Pandas' built-in time-based indexing to fetch future values without any loops. Here's how it works:
Ensure your index is properly formatted (datetime type and sorted):
# Convert index to datetime if it isn't already df.index = pd.to_datetime(df.index) # Sort the index (critical for fast reindexing) df = df.sort_index()Create an index of timestamps exactly 1 hour after each row's index:
future_timestamps = df.index + pd.Timedelta(hours=1)Fetch the future values and assign them as new columns:
# Reindex the original DataFrame to get values at the 1-hour-later timestamps future_values = df.reindex(future_timestamps)[["Glucosa", "Insulina", "Carbs"]] # Assign values to new columns in the original DataFrame df[["Glucosa1", "Insulina1", "Carbs1"]] = future_values.values
Why This Is So Much Faster
- This uses vectorized operations under the hood (implemented in C, not Python), which avoids the massive overhead of looping through each row in Python.
- Sorting the index lets Pandas use efficient binary search to locate future timestamps, rather than scanning the entire DataFrame for each row.
Handling Edge Cases
- If some timestamps don't have a corresponding entry exactly 1 hour later,
reindex()will fill those positions withNaN. You can customize this behavior with thefill_valueparameter:future_values = df.reindex(future_timestamps, fill_value=0)[["Glucosa", "Insulina", "Carbs"]] - If your time series has duplicate timestamps,
reindex()will return the first matching value by default. Adjust with themethodparameter if needed (e.g.,method="nearest"for the closest timestamp if exact matches are missing).
Performance Comparison
This approach should drop your runtime from ~8 seconds to well under 0.5 seconds, matching or exceeding the speed of optimized iterative functions. The exact speed depends on your DataFrame size, but it's night-and-day compared to row-by-row iteration.
内容的提问来源于stack exchange,提问作者Maganna Dev

