You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

时间序列不连续索引检测及连续段均值计算与新列生成问询

Hey there, let's work through your problem step by step. I'll start with generating column C as you described, then answer your two follow-up questions about detecting datetime index gaps and calculating means across continuous segments.

1. Generating Column C with Segment-Specific Averages

Based on your example, you need to calculate C(i) = A(i) + B(i) + average, where average is the mean of column B in the continuous segment before each gap. Let's break this into actionable code using pandas:

If you already know the exact segments (like your example with two intervals), you can manually calculate the averages and apply them to the relevant rows:

import pandas as pd

# Assume your DataFrame has a datetime index and columns 'A'/'B'
# Calculate averages for your specified segments
avg_segment1 = Data.loc['2005-03-02 02:05:30':'2005-03-02 02:07:30', 'B'].mean()
avg_segment2 = Data.loc['2005-03-02 02:08:30':'2005-03-02 02:11:00', 'B'].mean()

# Generate column C by applying the correct average to each row
Data['C'] = Data.apply(
    lambda row: row['A'] + row['B'] + (avg_segment1 if row.name <= '2005-03-02 02:07:30' else avg_segment2),
    axis=1
)

For a more scalable solution (if you have multiple gaps), we can automate segment detection (covered below) and map averages dynamically.

2. Detecting Gaps in Datetime Index

To find where your datetime index is discontinuous, compare the time difference between consecutive rows against your expected frequency (e.g., 30 seconds, inferred from your example timestamps):

# Calculate time differences between consecutive index entries
time_deltas = Data.index.to_series().diff()

# Define your expected interval (adjust this to match your data's frequency)
expected_interval = pd.Timedelta('30s')

# Create a mask where the gap exceeds the expected interval
gap_mask = time_deltas > expected_interval

# Get the positions right after each gap (or the missing times themselves)
gap_positions = Data.index[gap_mask]
missing_times = gap_positions - expected_interval

print("Missing datetime points:", missing_times.tolist())

This flags any rows where the previous entry was more than 30 seconds ago, indicating a gap in your time series.

3. Calculating Mean Over Continuous Index Segments

Once you've identified gaps, group your data into continuous segments and calculate the mean of B for each group:

# Assign a unique group ID to each continuous segment (increments at each gap)
Data['segment_group'] = gap_mask.cumsum()

# Calculate the mean of 'B' for each segment
segment_means = Data.groupby('segment_group')['B'].mean()

# Map the segment mean to each row and compute column C
Data['C'] = Data['A'] + Data['B'] + Data['segment_group'].map(segment_means)

# Clean up the temporary group column if needed
Data.drop('segment_group', axis=1, inplace=True)

This approach automatically handles any number of gaps, so you don't have to manually define segments each time.

Hope this helps you tackle your time series tasks smoothly!

内容的提问来源于stack exchange,提问作者dbose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:37:03