You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按日分组为每个变量计算线性回归斜率?

Solution to Extract Daily Linear Regression Slopes for Time-Series Variables

Got it, let's break this down clearly. You're right that TTR's ROC() isn't what you need here—it calculates point-to-point rate of change, not the overall linear trend slope from a regression. Here's a straightforward, reproducible approach using tidyverse tools in R:

Step 1: Prep Your Data

First, we need two key pieces of information from your Time column:

  • A date identifier to group data by day
  • A continuous numeric variable representing the time of day (e.g., minutes since midnight) to use as the independent variable in our regression.

We'll use lubridate for easy time manipulation:

# Load required packages
library(lubridate)
library(dplyr)
library(tidyr)

# Assume your dataset is named `df` with columns Time, Var1, Var2, Var3
df <- df %>%
  mutate(
    # Extract date for grouping
    date = as_date(Time),
    # Calculate minutes since midnight (continuous x variable for regression)
    minute_of_day = hour(Time) * 60 + minute(Time)
  )

Step 2: Calculate Daily Slopes (Two Approaches)

Approach 1: Long Format (Cleaner for Multiple Variables)

Convert your wide dataset to long format so we can process all Var columns in one go:

daily_slopes_long <- df %>%
  # Reshape to long format: one row per date-variable-observation
  pivot_longer(cols = starts_with("Var"), names_to = "variable", values_to = "value") %>%
  # Group by date and variable
  group_by(date, variable) %>%
  # Calculate slope, handle days with insufficient data (fewer than 2 points)
  summarize(
    slope = if(n() >= 2) {
      coef(lm(value ~ minute_of_day))[["minute_of_day"]]
    } else {
      NA_real_  # Assign NA if not enough data to run regression
    },
    .groups = "drop"
  )

This output will have columns date, variable, and slope—easy to filter, sort, or visualize.

Approach 2: Wide Format (Keep Original Variable Structure)

If you prefer to keep the output in wide format (one column per variable's slope), use this instead:

daily_slopes_wide <- df %>%
  group_by(date) %>%
  summarize(
    slope_var1 = if(n() >= 2) coef(lm(Var1 ~ minute_of_day))[["minute_of_day"]] else NA_real_,
    slope_var2 = if(n() >= 2) coef(lm(Var2 ~ minute_of_day))[["minute_of_day"]] else NA_real_,
    slope_var3 = if(n() >= 2) coef(lm(Var3 ~ minute_of_day))[["minute_of_day"]] else NA_real_,
    .groups = "drop"
  )

Key Notes

  • Why TTR::ROC() doesn't work: ROC() computes the percentage change between consecutive observations (e.g., (current - previous)/previous). This is a local, point-in-time change metric, not the global slope that describes the overall linear trend of the variable throughout the entire day.
  • Handling edge cases: The if(n() >=2) check prevents errors from trying to run a regression with too few data points. Adjust this threshold if you need stricter quality control (e.g., require at least 10 observations per day).
  • Alternative time variables: If you prefer, you can use decimal_date(Time) or another continuous time metric instead of minute_of_day—just ensure it’s a numeric variable that increases linearly throughout the day.

内容的提问来源于stack exchange,提问作者Kuo-Hsien Chang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:58:53