You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用线性回归预测DataFrame中不同Label对应的Value1未来值

多标签分组线性回归预测方案

需求说明

现有包含Label(含A、B、C三个唯一值)、Value1、Value2的DataFrame,需针对每个标签单独构建线性回归模型,预测当Value2=65000000时对应的Value1值。

示例DataFrame

import pandas as pd
data = {'Label': ['A','A','A','A','A','A','B','B','B','B','B','B','C','C','C','C','C','C'],
        'Value1': ['1672964520','1672966620','1672967460','1672969380','1672971840',
                   '1672972200','1672963800','1672966140', '1672967760','1672969020',
                   '1672970520', '1672971360','1672963200','1672964700','1672966260',
                   '1672967820', '1672969980', '1672971180'],
        'Value2': ['54727520', '54729380', '54740070', '54744720', '54775410', '54779130',
                   '59598560','59603190','59605060','59611320','59628900','59630950',
                   '58047810','58049680','58051550','58058460','58068740','58088280']}
df = pd.DataFrame(data)
print(df)

现有单标签实现(仅Label=A)

已实现单个标签A的预测逻辑,代码如下:

import numpy as np
import pandas as pd

data = {'Label': ['A','A','A','A','A','A',],
        'Value1': ['1672964520','1672966620','1672967460','1672969380','1672971840', '1672972200'],
        'Value2': ['54727520', '54729380', '54740070', '54744720', '54775410', '54779130']}

# 创建DataFrame
df = pd.DataFrame(data)
# 转换数据类型为整数
df["Value1"] = df.Value1.astype("int64")
df["Value2"] = df.Value2.astype("int")
# 计算均值
xmean = np.mean(df["Value1"])
ymean = np.mean(df["Value2"])
# 计算协方差和方差
df["xyCov"] = (df["Value1"] - xmean) * (df["Value2"] - ymean)
df["xVar"] = (df["Value1"] - xmean) ** 2
# 计算回归系数beta和截距alpha
beta = df["xyCov"].sum() / df["xVar"].sum()
alpha = ymean - (beta * xmean)
# 预测Value1值
Predicted_Value1 = (65000000 - alpha) / beta
print("Future A value", Predicted_Value1)

多标签适配方案

通过groupby按Label分组,对每个分组复用单标签的预测逻辑,即可实现多标签批量预测,完整代码如下:

import numpy as np
import pandas as pd

# 初始化数据
data = {'Label': ['A','A','A','A','A','A','B','B','B','B','B','B','C','C','C','C','C','C'],
        'Value1': ['1672964520','1672966620','1672967460','1672969380','1672971840',
                   '1672972200','1672963800','1672966140', '1672967760','1672969020',
                   '1672970520', '1672971360','1672963200','1672964700','1672966260',
                   '1672967820', '1672969980', '1672971180'],
        'Value2': ['54727520', '54729380', '54740070', '54744720', '54775410', '54779130',
                   '59598560','59603190','59605060','59611320','59628900','59630950',
                   '58047810','58049680','58051550','58058460','58068740','58088280']}
df = pd.DataFrame(data)

# 转换数据类型为整数
df["Value1"] = df["Value1"].astype("int64")
df["Value2"] = df["Value2"].astype("int")

# 定义每个分组的预测函数
def predict_value(group):
    x = group["Value1"]
    y = group["Value2"]
    xmean = np.mean(x)
    ymean = np.mean(y)
    # 计算beta和alpha
    beta = ((x - xmean) * (y - ymean)).sum() / ((x - xmean) ** 2).sum()
    alpha = ymean - beta * xmean
    # 预测Value1当Value2=65000000时的值
    predicted = (65000000 - alpha) / beta
    return predicted

# 按Label分组并应用预测函数
predictions = df.groupby("Label").apply(predict_value)

# 输出每个标签的预测结果
for label, pred in predictions.items():
    print(f"Predicted value of {label} = {pred:.2f}")

代码说明

  1. 数据类型转换:将Value1和Value2从字符串转为整数,确保后续数值计算正常。
  2. 分组处理:使用groupby("Label")将数据按标签拆分,每个组对应一个标签的独立数据集。
  3. 自定义预测函数:把单标签的线性回归计算逻辑封装成函数,每个分组独立计算回归参数并生成预测值。
  4. 结果输出:遍历分组后的预测结果,按要求格式打印每个标签对应的预测值。

内容的提问来源于stack exchange,提问作者user17236057

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 01:15:32