You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame中随机选取完整小时的4条15分钟数据并重复采样?

Efficiently Sample Full Hours of 15-Minute Data in Pandas

Got it, let's solve this problem without unnecessary overhead—since you want to avoid extra DataFrame queries, we'll lean into Pandas' vectorized operations and grouping logic to get the job done fast.

Step 1: Identify Hour Groups

First, we need a way to group all 15-minute records by their parent hour. If your StartTime column isn't already a datetime type, convert it first, then use dt.floor('H') to get the exact hour timestamp for each row. This gives us a shared identifier for all 4 records in the same hour.

import pandas as pd
import numpy as np

# Ensure StartTime is datetime (skip if already formatted)
df['StartTime'] = pd.to_datetime(df['StartTime'])

# Generate hour-level timestamps to group records (no permanent column needed)
hour_timestamps = df['StartTime'].dt.floor('H')

Step 2: Sample Random Hours

Next, grab all unique hour timestamps from the data, then randomly sample 1000 of them (allowing repeats with replace=True to match your 1000-sample requirement):

# Get all unique full hours present in the dataset
unique_hours = hour_timestamps.unique()

# Sample 1000 hours (with replacement to allow repeated hour selections)
sampled_hours = np.random.choice(unique_hours, size=1000, replace=True)

Step 3: Extract Sampled Hour Data

Now filter the original DataFrame to keep only rows belonging to our sampled hours. This uses a vectorized isin() check—way faster than looping through each sampled hour and running separate queries:

# Filter directly on the original DataFrame to get all 4 records per sampled hour
sampled_data = df[hour_timestamps.isin(sampled_hours)]

Optional: Condensed One-Liner

If you prefer to skip temporary variables and keep things concise (without modifying the original DataFrame), you can chain the operations:

sampled_data = df[
    df['StartTime'].dt.floor('H').isin(
        np.random.choice(df['StartTime'].dt.floor('H').unique(), size=1000, replace=True)
    )
]

Verify the Result

You can confirm the row count matches 4 * 1000 (assuming every hour has exactly 4 15-minute records) with:

print(len(sampled_data))  # Should output 4000

This approach keeps everything within the original DataFrame's context, uses fast vectorized operations instead of slow loops, and avoids redundant DataFrame queries—perfect for your performance needs.

内容的提问来源于stack exchange,提问作者jonasa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:59:21