You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为10分钟分箱的时间序列数据拟合最小二乘误差多项式函数?

时间序列10分钟分箱与多项式拟合实现

1. 数据预处理

先把time列转换成可处理的时间格式,确保后续分箱逻辑正常运行:

import pandas as pd
import numpy as np

# 加载数据(如果是本地CSV文件,替换为实际路径)
df = pd.read_csv('your_data.csv')
# 转换time列为datetime类型(仅时间格式也可正常处理)
df['time'] = pd.to_datetime(df['time'], format='%H:%M')

2. 按10分钟间隔分箱

使用pd.Grouper按10分钟粒度分组,确保分箱区间匹配你要求的模式(如15:20-15:30、15:30-15:40):

# 按10分钟分箱,origin='start_day'确保分箱从当日00:00开始按10分钟划分
bin_groups = df.groupby(pd.Grouper(key='time', freq='10min', origin='start_day'))
# 为每条数据标记所属分箱编号(可选,方便后续查看)
df['bin_id'] = bin_groups.ngroup()

3. 分箱内最小二乘多项式拟合

定义拟合函数,对每个分箱的数据用最小二乘法拟合指定次数的多项式:

def fit_poly(group, degree=2):
    # 将分箱内的时间转换为相对于分箱起始点的分钟数,简化拟合计算
    group['relative_minutes'] = (group['time'] - group['time'].min()).dt.total_seconds() / 60
    # 最小二乘拟合多项式,返回系数(从高次项到常数项)
    coeffs = np.polyfit(group['relative_minutes'], group['var'], deg=degree)
    # 计算拟合值
    group['fitted_var'] = np.polyval(coeffs, group['relative_minutes'])
    # 保存多项式系数(可选)
    group['poly_coeffs'] = [coeffs] * len(group)
    return group

# 应用拟合函数,这里指定2次多项式,可根据需求改为1(线性)、3次等
df_fitted = bin_groups.apply(fit_poly, degree=2).reset_index(drop=True)

4. 结果验证与优化

  • 查看指定分箱的原始值与拟合值:
# 查看第一个分箱的数据
print(df_fitted[df_fitted['bin_id'] == 0][['time', 'var', 'fitted_var']])
  • 过滤数据点不足的分箱(避免因数据量太少导致拟合报错):
degree = 2
# 仅保留数据点数量≥多项式次数+1的分箱
valid_groups = bin_groups.filter(lambda x: len(x) >= degree + 1)
df_fitted = valid_groups.apply(fit_poly, degree=degree).reset_index(drop=True)

内容的提问来源于stack exchange,提问作者prof31

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:30:40