You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按ID计算非等距日期的7天移动平均与移动斜率

7-Day Moving Average & Slope for Grouped Uneven Time Series

Hey there! Let's work through how to calculate the 7-day moving average and 7-day moving slope for oldvar across your per-id time series data. Since your data has uneven dates, missing entries, and variable observation counts per id, we need solutions that handle grouped data and date-based windows (not just fixed row counts). Here's how to do this in both R and Python, which are the most common tools for this kind of task:

Key Context to Note

First, let's clarify: we're defining a 7-day moving window as all observations from the current date going back exactly 7 days (not a fixed number of rows). This is critical because your dates aren't evenly spaced.


R Implementation

We'll use dplyr for grouping, slider for flexible date-based window calculations, and lubridate to handle dates.

Step 1: Install Required Packages

install.packages(c("dplyr", "slider", "lubridate"))

Step 2: Full Code

Assuming your data frame is named df with columns id, date, and oldvar:

library(dplyr)
library(slider)
library(lubridate)

# Clean and process the data
df_clean <- df %>%
  # Ensure date is parsed as a Date type
  mutate(date = ymd(date)) %>%
  # Group by id and sort observations by date within each group
  group_by(id) %>%
  arrange(date, .by_group = TRUE) %>%
  # Calculate 7-day moving average
  mutate(
    ma_7d = slide_dbl(
      .x = oldvar,
      .i = date,
      .f = ~mean(.x, na.rm = TRUE),
      .before = days(7),  # Window = current date minus 7 days
      .complete = FALSE   # Allow partial windows (e.g., first few observations)
    ),
    # Calculate 7-day moving slope (requires at least 2 points in window)
    slope_7d = slide_dbl(
      .x = tibble(date_num = as.numeric(date), var = oldvar),
      .i = date,
      .f = ~{
        if(nrow(.x) >= 2){
          # Fit linear regression and extract slope
          lm(var ~ date_num, data = .x)$coefficients[["date_num"]]
        } else {
          NA_real_  # Return NA if not enough points
        }
      },
      .before = days(7),
      .complete = FALSE
    )
  ) %>%
  ungroup()

Explanation

  • .before = days(7) ensures we only include observations from the past 7 days relative to the current row's date.
  • Set .complete = TRUE if you want to only calculate stats for windows that have at least one observation from every day in the 7-day range (not recommended for uneven dates).
  • na.rm = TRUE skips any missing oldvar values when calculating the average.

Python Implementation

We'll use pandas for grouping and time window handling, plus scipy.stats to calculate linear regression slopes.

Step 1: Install Required Packages

pip install pandas numpy scipy

Step 2: Full Code

Assuming your data frame is named df with columns id, date, and oldvar:

import pandas as pd
import numpy as np
from scipy.stats import linregress

# Parse date column to datetime type
df['date'] = pd.to_datetime(df['date'])

def compute_7d_metrics(group):
    # Sort group by date first
    group_sorted = group.sort_values('date').reset_index(drop=True)
    
    # Calculate 7-day moving average using time-based rolling window
    group_sorted['ma_7d'] = group_sorted['oldvar'].rolling(
        window='7D',
        on='date',
        closed='right',  # Include the current date in the window
        skipna=True
    ).mean()
    
    # Calculate 7-day moving slope
    slopes = []
    for idx, row in group_sorted.iterrows():
        # Define window range: current date minus 7 days to current date
        window_end = row['date']
        window_start = window_end - pd.Timedelta(days=7)
        window_data = group_sorted[
            (group_sorted['date'] >= window_start) & 
            (group_sorted['date'] <= window_end)
        ]
        
        # Only calculate slope if we have at least 2 data points
        if len(window_data) >= 2:
            # Convert dates to numeric (days since the first date in the window)
            date_numeric = (window_data['date'] - window_data['date'].min()).dt.days
            slope, _, _, _, _ = linregress(date_numeric, window_data['oldvar'])
            slopes.append(slope)
        else:
            slopes.append(np.nan)
    
    group_sorted['slope_7d'] = slopes
    return group_sorted

# Apply the function to each id group
df_clean = df.groupby('id').apply(compute_7d_metrics).reset_index(drop=True)

Explanation

  • rolling(window='7D', on='date') creates a time-based window instead of a fixed row count, which works for uneven dates.
  • closed='right' includes the current row's date in the window; use closed='left' if you want to exclude the current date.
  • We convert dates to numeric values for the linear regression because linregress can't work directly with datetime objects.

Important Notes

  • Always validate that your date column is properly parsed as a date/datetime type—this is critical for accurate window calculations.
  • If you have missing oldvar values, adjust the na.rm (R) or skipna (Python) parameters to handle them as needed.
  • For edge cases (e.g., an id with only 3 observations), the slope will return NA for the first row (since you need at least 2 points to calculate a slope).

内容的提问来源于stack exchange,提问作者syork

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:44:26