如何从DataFrame中随机选取完整小时的4条15分钟数据并重复采样?
Got it, let's solve this problem without unnecessary overhead—since you want to avoid extra DataFrame queries, we'll lean into Pandas' vectorized operations and grouping logic to get the job done fast.
Step 1: Identify Hour Groups
First, we need a way to group all 15-minute records by their parent hour. If your StartTime column isn't already a datetime type, convert it first, then use dt.floor('H') to get the exact hour timestamp for each row. This gives us a shared identifier for all 4 records in the same hour.
import pandas as pd import numpy as np # Ensure StartTime is datetime (skip if already formatted) df['StartTime'] = pd.to_datetime(df['StartTime']) # Generate hour-level timestamps to group records (no permanent column needed) hour_timestamps = df['StartTime'].dt.floor('H')
Step 2: Sample Random Hours
Next, grab all unique hour timestamps from the data, then randomly sample 1000 of them (allowing repeats with replace=True to match your 1000-sample requirement):
# Get all unique full hours present in the dataset unique_hours = hour_timestamps.unique() # Sample 1000 hours (with replacement to allow repeated hour selections) sampled_hours = np.random.choice(unique_hours, size=1000, replace=True)
Step 3: Extract Sampled Hour Data
Now filter the original DataFrame to keep only rows belonging to our sampled hours. This uses a vectorized isin() check—way faster than looping through each sampled hour and running separate queries:
# Filter directly on the original DataFrame to get all 4 records per sampled hour sampled_data = df[hour_timestamps.isin(sampled_hours)]
Optional: Condensed One-Liner
If you prefer to skip temporary variables and keep things concise (without modifying the original DataFrame), you can chain the operations:
sampled_data = df[ df['StartTime'].dt.floor('H').isin( np.random.choice(df['StartTime'].dt.floor('H').unique(), size=1000, replace=True) ) ]
Verify the Result
You can confirm the row count matches 4 * 1000 (assuming every hour has exactly 4 15-minute records) with:
print(len(sampled_data)) # Should output 4000
This approach keeps everything within the original DataFrame's context, uses fast vectorized operations instead of slow loops, and avoids redundant DataFrame queries—perfect for your performance needs.
内容的提问来源于stack exchange,提问作者jonasa

