使用pd.IntervalIndex匹配时间区间关联Action字段报错排查
Let's break down what's going wrong and how to fix it:
Why You're Getting the Error
Your code is throwing that TypeError because you're not creating the IntervalIndex correctly. When you pass two Series directly to the IntervalIndex constructor, it doesn't interpret them as paired start/end times for intervals — instead, it tries to treat each individual timestamp as an interval, which obviously doesn't work.
Corrected Code Using IntervalIndex
The right way to create your time-based interval index is with IntervalIndex.from_arrays(), which explicitly pairs your start and end times into valid intervals:
import pandas as pd # Create the IntervalIndex properly by pairing start and end times idx = pd.IntervalIndex.from_arrays( B['Video Start Time'], B['Video End Time'], closed='both' ) # Get the position of each Timestamp in A within the interval index match_indices = idx.get_indexer(A['Timestamp']) # Map the matching intervals to their corresponding Action values A['event'] = B.loc[match_indices, 'Action'].values # Handle cases where a Timestamp doesn't fall into any interval (optional) A['event'] = A['event'].fillna('No Match')
How This Works
from_arrays()ensures each row in B'sVideo Start Timeis paired with the correspondingVideo End Timeto form a valid datetime interval.get_indexer()then finds which interval each timestamp in A belongs to (returning-1if there's no match).- We use those indices to pull the correct
Actionvalues from B and assign them to A's neweventcolumn.
Testing this with your sample data: All timestamps in A fall within the first interval in B, so every row in A's event column will be set to Relaxation — which is exactly what you'd expect.
Alternative Approach: Using merge_asof
If you prefer a more intuitive method for time-based matching, pd.merge_asof is a great alternative. It's designed for associating rows between DataFrames based on time proximity, and we can add a filter to ensure timestamps fall within the full interval:
# Sort both DataFrames (required for merge_asof) A_sorted = A.sort_values('Timestamp') B_sorted = B.sort_values('Video Start Time') # Perform the as-of merge to match timestamps to the most recent start time merged = pd.merge_asof( A_sorted, B_sorted, left_on='Timestamp', right_on='Video Start Time', direction='backward' ) # Filter out any rows where the timestamp exceeds the video end time merged = merged[merged['Timestamp'] <= merged['Video End Time']] # Merge the results back to the original A DataFrame (preserves original order) A = A.merge(merged[['Timestamp', 'Action']], on='Timestamp', how='left') A.rename(columns={'Action': 'event'}, inplace=True) A['event'] = A['event'].fillna('No Match')
This method avoids manual interval creation and is often more efficient with large datasets.
内容的提问来源于stack exchange,提问作者Sundar N

