基于Pandas处理分时股票数据:合并9:00-9:16数据至9:16节点
Hey there! Let's work through fixing this stock data processing problem together. The issue with your current code is that it treats all 9:00-9:16 data across every trading day as a single group, which leads to incorrect global values instead of per-day calculations. Let's fix this step by step.
Your original code pulls all 9:00-9:16 data into one batch and uses global values (like the first open across all days, or overall max high) to update every 9:16 row. But we need to calculate these values per individual trading day since each day's pre-9:16 data is independent.
First, make sure your DataFrame's index is a datetime type (if it isn't already):
import pandas as pd # Convert index to datetime if needed df.index = pd.to_datetime(df.index)
Then use this code to process the data correctly:
# 1. Keep all data from 9:16 onwards (including the 9:16 timestamp itself) post_916_data = df[df.index.time >= pd.to_datetime("09:16").time()] # 2. Extract 9:00-9:16 data and calculate required values per trading day pre_916_data = df.between_time("09:00", "09:16") daily_pre_calculations = pre_916_data.groupby(pre_916_data.index.date).agg( open=("open", "first"), # First open value of the 9:00-9:16 window high=("high", "max"), # Highest high in the window low=("low", "min"), # Lowest low in the window # Grab the close value exactly at 9:16 close=("close", lambda x: x[x.index.time == pd.to_datetime("09:16").time()].iloc[0]) ) # 3. Update the 9:16 row in the post-916 data with per-day calculations for trade_date, calc_row in daily_pre_calculations.iterrows(): target_timestamp = pd.to_datetime(f"{trade_date} 09:16:00") if target_timestamp in post_916_data.index: post_916_data.loc[target_timestamp, ["open", "high", "low", "close"]] = calc_row.values # 4. Final processed data (sorted to maintain time order) final_processed_df = post_916_data.sort_index()
- Isolate post-9:16 data: We start by keeping all data from 9:16 onwards—this is our base, and we only need to modify the 9:16 rows.
- Per-day pre-9:16 calculations: Grouping by trading date ensures we compute values (first open, max high, min low, 9:16 close) separately for each day.
- Update 9:16 rows: We loop through each day's calculations and overwrite the corresponding 9:16 row in our base data.
- Sort for consistency: Ensures the final DataFrame maintains a chronological order.
Using your sample data:
- Original 9:16 row:
116.10 117.80 117.00 113.00 - After processing, this row becomes:
116.00 117.80 116.00 113.00
Which matches your requirements: open from the first pre-9:16 entry, max high from the window, min low from the window, and close from the 9:16 timestamp.
内容的提问来源于stack exchange,提问作者amitnair92

