You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在StatsModels中为面板相似变量构建Local与TVP-VAR模型

批量处理面板数据的Local模型(StatsModels)

针对你提到的宽格式面板数据(时间为索引,列名含国家代码+变量类型),可以通过变量分组+函数封装+批量迭代的方式实现同类变量的统一处理,以下是具体实现方案:


一、按变量类型归类列名

首先基于列名的命名规则(国家代码_变量类型),用字典将同类变量(如所有通胀、GDP变量)分组:

import pandas as pd
import numpy as np
import statsmodels.api as sm
import matplotlib.pyplot as plt

# 假设你的数据框为dta,已设置时间为索引
# 提取所有唯一的变量类型
var_types = list(set(col.split('_')[-1] for col in dta.columns))

# 创建变量分组字典:键为变量类型,值为对应所有国家的列名列表
var_groups = {}
for var_type in var_types:
    var_groups[var_type] = [col for col in dta.columns if col.endswith(f'_{var_type}')]

# 示例:var_groups['inflation'] 将包含AUS_inflation、USA_inflation等所有通胀列

二、封装Local模型运行函数

将你提供的单变量模型代码封装为可复用函数,支持输入时间序列、输出模型结果与模拟数据:

def run_local_level_model(ts_data, var_label):
    """
    运行Local Level模型,返回拟合结果与模拟状态数据
    参数:
        ts_data: pandas.Series,目标时间序列
        var_label: str,变量标识(如"AUS_gdp",用于图表标注)
    返回:
        res: 模型拟合结果对象
        simulated_kfs: KFS模拟状态数据(DataFrame,含20次模拟列)
        simulated_cfa: CFA模拟状态数据(DataFrame,含20次模拟列)
    """
    # 构建Local Level模型
    mod = sm.tsa.UnobservedComponents(ts_data, 'llevel')
    # 极大似然拟合参数
    res = mod.fit()
    print(f"{var_label} 模型参数:\n{res.params}\n")
    
    # 初始化模拟平滑器
    sim_kfs = mod.simulation_smoother()
    sim_cfa = mod.simulation_smoother(method='cfa')
    
    nsimulations = 20
    # 创建模拟数据存储容器
    simulated_kfs = pd.DataFrame(
        np.zeros((mod.nobs, nsimulations)), index=ts_data.index
    )
    simulated_cfa = pd.DataFrame(
        np.zeros((mod.nobs, nsimulations)), index=ts_data.index
    )
    
    # 执行模拟
    for i in range(nsimulations):
        sim_kfs.simulate()
        simulated_kfs.iloc[:, i] = sim_kfs.simulated_state[0]
        
        sim_cfa.simulate()
        simulated_cfa.iloc[:, i] = sim_cfa.simulated_state[0]
    
    # 绘制KFS/CFA模拟对比图
    fig, axes = plt.subplots(2, figsize=(15, 6))
    # KFS模拟结果
    ts_data.plot(ax=axes[0], color='k')
    axes[0].set_title(f"{var_label} - KFS方法模拟(MLE参数)")
    simulated_kfs.plot(ax=axes[0], color='C0', alpha=0.25, legend=False)
    # CFA模拟结果
    ts_data.plot(ax=axes[1], color='k')
    axes[1].set_title(f"{var_label} - CFA方法模拟(MLE参数)")
    simulated_cfa.plot(ax=axes[1], color='C0', alpha=0.25, legend=False)
    
    handles, labels = axes[0].get_legend_handles_labels()
    axes[0].legend(handles[:2], ['原始数据', '模拟状态'])
    fig.tight_layout()
    
    # 绘制替代参数化的平滑趋势模拟图
    fig, ax = plt.subplots(figsize=(15, 3))
    mod.update([4, 0.05])
    for i in range(nsimulations):
        sim_kfs.simulate()
        ax.plot(ts_data.index, sim_kfs.simulated_state[0], color='C0', alpha=0.25)
    ts_data.plot(ax=ax, color='k', zorder=-1)
    ax.set_title(f"{var_label} - 平滑趋势的替代参数化模拟")
    ax.legend(['模拟状态', '原始数据'])
    fig.tight_layout()
    
    return res, simulated_kfs, simulated_cfa

三、批量迭代所有变量并存储结果

用嵌套字典存储所有模型结果,方便后续调用分析:

# 初始化结果存储容器:外层键为变量类型,内层键为国家代码
all_results = {}

for var_type, cols in var_groups.items():
    all_results[var_type] = {}
    for col in cols:
        # 拆分国家代码与变量类型
        country_code = col.split('_')[0]
        # 获取当前列的时间序列数据
        ts_data = dta[col]
        # 运行模型并存储结果
        res, sim_kfs, sim_cfa = run_local_level_model(ts_data, col)
        all_results[var_type][country_code] = {
            'model_result': res,
            'simulated_kfs': sim_kfs,
            'simulated_cfa': sim_cfa
        }

四、转换为经典面板格式结果

如果需要将模拟数据转换为经典面板结构(时间、国家、变量、模拟值),可执行以下代码:

# 整理KFS模拟结果为面板格式
kfs_panel_list = []
for var_type, country_dict in all_results.items():
    for country_code, result_dict in country_dict.items():
        sim_df = result_dict['simulated_kfs']
        # 将宽格式模拟数据转为长格式
        long_df = sim_df.melt(ignore_index=False, var_name='simulation_id', value_name='simulated_value')
        long_df['country'] = country_code
        long_df['variable_type'] = var_type
        kfs_panel_list.append(long_df.reset_index())

# 合并为最终面板数据框
kfs_panel_df = pd.concat(kfs_panel_list, ignore_index=True)
# 最终列:时间索引列、simulation_id、simulated_value、country、variable_type

扩展说明

  • 命名规则适配:如果你的列名命名规则不同(如变量类型在前),只需调整split('_')的索引位置即可。
  • TVP-VAR适配:后续运行TVP-VAR模型时,可复用相同的变量分组逻辑,先按类型整理数据,再批量构建模型。
  • 结果复用:存储在all_results中的模型结果可直接用于参数分析、预测等后续操作,无需重复拟合。

内容的提问来源于stack exchange,提问作者John D.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 19:50:17