Python中对非均匀时间序列上采样并保留原始点值
Upsampling Uneven Time Series to 50 Points (Preserving Original Data)
Got it, let's solve this problem step by step. The core requirements are:
- Upsample your 7-point uneven time series to 50 total points
- Keep all original data points
- Add more new points to longer time intervals (since they have more "space" for samples)
Approach Overview
- Convert string timestamps to datetime objects (so we can calculate time differences)
- Calculate how many additional points we need (50 - 7 = 43)
- Allocate these 43 points proportionally to each interval based on its length (longer intervals get more points)
- Generate new timestamps in each interval, then merge, sort, and deduplicate to get the final list
Solution 1: Using Python Standard Library
This method uses only built-in modules, no external dependencies:
from datetime import datetime, timedelta import math # Original time series ls = ['2016-01-30 12:10:00', '2016-01-30 12:23:35', '2016-01-30 12:24:14', '2016-01-30 12:24:51', '2016-01-30 12:25:00', '2016-01-30 12:26:49', '2016-01-30 12:27:36'] # Step 1: Convert strings to datetime objects dt_list = [datetime.strptime(t, '%Y-%m-%d %H:%M:%S') for t in ls] total_needed = 50 additional_points = total_needed - len(dt_list) # 43 points to add # Step 2: Calculate time intervals between consecutive points deltas = [dt_list[i+1] - dt_list[i] for i in range(len(dt_list)-1)] total_duration = dt_list[-1] - dt_list[0] # Step 3: Allocate points proportionally to each interval points_per_interval = [] for delta in deltas: ratio = delta.total_seconds() / total_duration.total_seconds() # Use floor to get initial count, then adjust for remaining points num_points = math.floor(ratio * additional_points) points_per_interval.append(num_points) # Adjust to ensure total additional points sum to 43 remaining = additional_points - sum(points_per_interval) # Distribute remaining points to the longest intervals first sorted_interval_indices = sorted(range(len(deltas)), key=lambda x: deltas[x], reverse=True) for i in range(remaining): points_per_interval[sorted_interval_indices[i]] += 1 # Step 4: Generate new timestamps and merge with original new_dt_list = dt_list.copy() for i in range(len(deltas)): start, end = dt_list[i], dt_list[i+1] num_points = points_per_interval[i] if num_points <= 0: continue # Calculate step size to insert points evenly between start and end step = (end - start) / (num_points + 1) # +1 avoids overlapping with original points for j in range(1, num_points + 1): new_time = start + step * j new_dt_list.append(new_time) # Sort and remove duplicates (just in case of rounding edge cases) new_dt_list = sorted(list(set(new_dt_list))) # Convert back to string format final_time_list = [t.strftime('%Y-%m-%d %H:%M:%S') for t in new_dt_list] # Verify the result print(f"Final total points: {len(final_time_list)}") # Output: 50
Solution 2: Using Pandas (Simpler for Time Series)
If you're working with time series regularly, Pandas makes this task more concise:
import pandas as pd ls = ['2016-01-30 12:10:00', '2016-01-30 12:23:35', '2016-01-30 12:24:14', '2016-01-30 12:24:51', '2016-01-30 12:25:00', '2016-01-30 12:26:49', '2016-01-30 12:27:36'] # Create a Series with datetime index (values don't matter here) ts = pd.Series(range(len(ls)), index=pd.to_datetime(ls)) total_needed = 50 additional_points = total_needed - len(ts) # Calculate intervals and proportional point allocation intervals = ts.index[1:] - ts.index[:-1] total_duration = ts.index[-1] - ts.index[0] ratios = intervals / total_duration points_per_interval = (ratios * additional_points).astype(int) # Adjust for remaining points remaining = additional_points - points_per_interval.sum() largest_intervals = intervals.argsort()[::-1][:remaining] points_per_interval.iloc[largest_intervals] += 1 # Generate new timestamps new_indices = ts.index.copy() for i in range(len(intervals)): start, end = ts.index[i], ts.index[i+1] num_points = points_per_interval.iloc[i] if num_points <= 0: continue # Generate evenly spaced points between start/end (exclude original points) new_times = pd.date_range(start=start, end=end, periods=num_points + 2)[1:-1] new_indices = new_indices.append(new_times) # Sort and deduplicate new_indices = new_indices.sort_values().drop_duplicates() # Convert back to string format final_list = new_indices.strftime('%Y-%m-%d %H:%M:%S').tolist() print(f"Final total points: {len(final_list)}") # Output: 50
Key Notes
- Both methods ensure original points are preserved (we never overwrite or remove them)
- Longer intervals get more new points because we allocate based on the interval's proportion of the total duration
- The new points are evenly spaced within each interval (you can adjust this if you want non-uniform spacing, but the problem states position isn't strictly required)
内容的提问来源于stack exchange,提问作者mastersom
相关产品推荐
相关产品推荐

