You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中对非均匀时间序列上采样并保留原始点值

Upsampling Uneven Time Series to 50 Points (Preserving Original Data)

Got it, let's solve this problem step by step. The core requirements are:

  • Upsample your 7-point uneven time series to 50 total points
  • Keep all original data points
  • Add more new points to longer time intervals (since they have more "space" for samples)

Approach Overview

  1. Convert string timestamps to datetime objects (so we can calculate time differences)
  2. Calculate how many additional points we need (50 - 7 = 43)
  3. Allocate these 43 points proportionally to each interval based on its length (longer intervals get more points)
  4. Generate new timestamps in each interval, then merge, sort, and deduplicate to get the final list

Solution 1: Using Python Standard Library

This method uses only built-in modules, no external dependencies:

from datetime import datetime, timedelta
import math

# Original time series
ls = ['2016-01-30 12:10:00', '2016-01-30 12:23:35', '2016-01-30 12:24:14', 
      '2016-01-30 12:24:51', '2016-01-30 12:25:00', '2016-01-30 12:26:49', 
      '2016-01-30 12:27:36']

# Step 1: Convert strings to datetime objects
dt_list = [datetime.strptime(t, '%Y-%m-%d %H:%M:%S') for t in ls]
total_needed = 50
additional_points = total_needed - len(dt_list)  # 43 points to add

# Step 2: Calculate time intervals between consecutive points
deltas = [dt_list[i+1] - dt_list[i] for i in range(len(dt_list)-1)]
total_duration = dt_list[-1] - dt_list[0]

# Step 3: Allocate points proportionally to each interval
points_per_interval = []
for delta in deltas:
    ratio = delta.total_seconds() / total_duration.total_seconds()
    # Use floor to get initial count, then adjust for remaining points
    num_points = math.floor(ratio * additional_points)
    points_per_interval.append(num_points)

# Adjust to ensure total additional points sum to 43
remaining = additional_points - sum(points_per_interval)
# Distribute remaining points to the longest intervals first
sorted_interval_indices = sorted(range(len(deltas)), key=lambda x: deltas[x], reverse=True)
for i in range(remaining):
    points_per_interval[sorted_interval_indices[i]] += 1

# Step 4: Generate new timestamps and merge with original
new_dt_list = dt_list.copy()
for i in range(len(deltas)):
    start, end = dt_list[i], dt_list[i+1]
    num_points = points_per_interval[i]
    if num_points <= 0:
        continue
    # Calculate step size to insert points evenly between start and end
    step = (end - start) / (num_points + 1)  # +1 avoids overlapping with original points
    for j in range(1, num_points + 1):
        new_time = start + step * j
        new_dt_list.append(new_time)

# Sort and remove duplicates (just in case of rounding edge cases)
new_dt_list = sorted(list(set(new_dt_list)))

# Convert back to string format
final_time_list = [t.strftime('%Y-%m-%d %H:%M:%S') for t in new_dt_list]

# Verify the result
print(f"Final total points: {len(final_time_list)}")  # Output: 50

Solution 2: Using Pandas (Simpler for Time Series)

If you're working with time series regularly, Pandas makes this task more concise:

import pandas as pd

ls = ['2016-01-30 12:10:00', '2016-01-30 12:23:35', '2016-01-30 12:24:14', 
      '2016-01-30 12:24:51', '2016-01-30 12:25:00', '2016-01-30 12:26:49', 
      '2016-01-30 12:27:36']

# Create a Series with datetime index (values don't matter here)
ts = pd.Series(range(len(ls)), index=pd.to_datetime(ls))
total_needed = 50
additional_points = total_needed - len(ts)

# Calculate intervals and proportional point allocation
intervals = ts.index[1:] - ts.index[:-1]
total_duration = ts.index[-1] - ts.index[0]
ratios = intervals / total_duration
points_per_interval = (ratios * additional_points).astype(int)

# Adjust for remaining points
remaining = additional_points - points_per_interval.sum()
largest_intervals = intervals.argsort()[::-1][:remaining]
points_per_interval.iloc[largest_intervals] += 1

# Generate new timestamps
new_indices = ts.index.copy()
for i in range(len(intervals)):
    start, end = ts.index[i], ts.index[i+1]
    num_points = points_per_interval.iloc[i]
    if num_points <= 0:
        continue
    # Generate evenly spaced points between start/end (exclude original points)
    new_times = pd.date_range(start=start, end=end, periods=num_points + 2)[1:-1]
    new_indices = new_indices.append(new_times)

# Sort and deduplicate
new_indices = new_indices.sort_values().drop_duplicates()

# Convert back to string format
final_list = new_indices.strftime('%Y-%m-%d %H:%M:%S').tolist()

print(f"Final total points: {len(final_list)}")  # Output: 50

Key Notes

  • Both methods ensure original points are preserved (we never overwrite or remove them)
  • Longer intervals get more new points because we allocate based on the interval's proportion of the total duration
  • The new points are evenly spaced within each interval (you can adjust this if you want non-uniform spacing, but the problem states position isn't strictly required)

内容的提问来源于stack exchange,提问作者mastersom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:27:55