DataFrame时间戳筛选遇TypeError,求分段方案及范围判断局限性
Hey there! Let's break down your two technical questions step by step, especially since you're working with a sizeable 1.1M+ row dataset and trying to split it into 10-second chunks.
Absolutely—here are the key gotchas to watch out for:
- Data type mismatches: If your
Global Timecolumn isn’t stored as a numeric type (likeint64orfloat64), direct comparisons with integer timestamps will fail. For example, if it’s stored as strings, comparing strings to integers will throw errors or return incorrect results. - Time zone confusion: If your timestamps are in UTC but your data was collected in a local time zone, your range filters might include/exclude the wrong records. Always confirm the time zone context of your timestamps before filtering.
- Boundary edge cases: Depending on how you define your ranges (inclusive vs. exclusive), you might accidentally exclude records that fall exactly on the boundary. Your current syntax (
>=and<=) is inclusive, which is fine—but just be intentional about it. - Performance on large datasets: Without an index on
Global Time, filtering 1M+ rows can be slow. Adding an index withsmall_df = small_df.set_index('Global Time')will speed up range queries significantly.
The error TypeError: <class 'int'> type object 1118846979200 usually points to a data type mismatch between your timestamp values and the Global Time column. Here's how to fix it:
First, check the data type of Global Time
Run this to confirm what type your column is:
print(small_df['Global Time'].dtype)
If it's stored as strings (object dtype)
Convert it to integers first—handle any messy data if needed:
# Try direct conversion to int small_df['Global Time'] = small_df['Global Time'].astype(int) # If there are non-numeric values, use error-safe conversion small_df['Global Time'] = pd.to_numeric(small_df['Global Time'], errors='coerce') # Drop rows with NaN values from failed conversions small_df = small_df.dropna(subset=['Global Time'])
Alternative filtering syntax (avoids parentheses issues)
Sometimes nested boolean conditions can cause unexpected errors—use query() for cleaner, more reliable code:
start_ts = 1118846979200 end_ts = 1118846989200 filtered_df = small_df.query('`Global Time` >= @start_ts and `Global Time` <= @end_ts')
Automate 10-second chunking
Instead of manually writing ranges for each window, use grouping to split your DataFrame automatically (way more efficient for large datasets):
# Assuming timestamps are in milliseconds (10 seconds = 10000 ms) # Calculate the start timestamp of each 10-second chunk small_df['10s_chunk_start'] = (small_df['Global Time'] // 10000) * 10000 # Group by chunk start time to get all subsets chunked_dfs = {chunk_start: group for chunk_start, group in small_df.groupby('10s_chunk_start')} # Access a specific chunk like this: target_chunk = chunked_dfs[1118846979200]
内容的提问来源于stack exchange,提问作者Angai

