基于宏观时间序列的微观时间序列缺失值插值填充及约束条件设置技术问询
Awesome question! Let's walk through how to fill those missing micro-data values using your macro data while matching its dynamics, plus how to add those custom constraints you're asking about.
Since your micro data (col1, col2) likely co-moves with the macro data (col_macro), we can leverage the macro series' change rates and shocks to generate realistic interpolations. Below are practical implementations using Python (Pandas/Numpy/SciPy):
First, let's load your sample data:
import pandas as pd import numpy as np # Your input data data = { "col1": [1, np.nan, np.nan, np.nan, np.nan, 8], "col2": [2, np.nan, np.nan, np.nan, np.nan, 2], "col_macro": [3, 4, 7, 7, 13, 18] } df = pd.DataFrame(data)
We have two reliable approaches here:
Approach 1: Macro Change Rate-Driven Interpolation
This method uses the macro series' percentage changes to drive micro data interpolation, then scales the result to match the known start/end values:
def interpolate_with_macro(df, micro_col, macro_col): start_val = df[micro_col].iloc[0] end_val = df[micro_col].iloc[-1] # Get macro changes between known micro points macro_changes = df[macro_col].pct_change().iloc[1:-1] # Build initial interpolation using macro changes interpolated = [start_val] for change in macro_changes: interpolated.append(interpolated[-1] * (1 + change)) interpolated.append(end_val) # Scale the penultimate value to ensure exact match with end_val scaling_factor = end_val / interpolated[-2] interpolated[-2] *= scaling_factor df[f"{micro_col}_macro_interp"] = interpolated return df # Apply to your columns df = interpolate_with_macro(df, "col1", "col_macro") df = interpolate_with_macro(df, "col2", "col_macro")
This ensures your filled micro data follows the same ups/downs as col_macro while respecting the known start/end points.
Approach 2: Regression-Based Interpolation
If your micro and macro data have a statistical relationship, train a simple regression model on non-missing micro values, then predict the gaps:
from sklearn.linear_model import LinearRegression def reg_interpolate(df, micro_col, macro_col): # Isolate non-missing data non_missing = df.dropna(subset=[micro_col]) X = non_missing[[macro_col]] y = non_missing[micro_col] # Train model model = LinearRegression() model.fit(X, y) # Predict missing values df[f"{micro_col}_reg_interp"] = model.predict(df[[macro_col]]) return df df = reg_interpolate(df, "col1", "col_macro")
This method keeps micro data aligned with macro trends in a statistically consistent way.
Absolutely, you can enforce constraints like constant values (when endpoints are similar) or strict monotonicity:
Constraint 1: Constant Interpolation When Endpoints Are Similar
If the start and end values of a micro column are nearly identical, fill all gaps with that value:
def interpolate_with_constant_constraint(df, micro_col, macro_col, threshold=0.01): start_val = df[micro_col].iloc[0] end_val = df[micro_col].iloc[-1] # Check if endpoints are close enough if abs(start_val - end_val) < threshold: df[f"{micro_col}_const_constant"] = start_val return df # Fall back to macro-driven interpolation if not return interpolate_with_macro(df, micro_col, macro_col) df = interpolate_with_constant_constraint(df, "col2", "col_macro")
Constraint 2: Strictly Increasing/Decreasing Interpolation
Force the filled sequence to never reverse direction using SciPy's interpolation and post-processing:
from scipy.interpolate import interp1d def interpolate_monotonic(df, micro_col, macro_col, increasing=True): # Get indices and values of non-missing micro data non_missing_idx = df[micro_col].dropna().index non_missing_vals = df[micro_col].dropna().values macro_vals = df[macro_col].values # Build interpolation function tied to macro data f = interp1d(df.loc[non_missing_idx, macro_col], non_missing_vals, kind="linear", fill_value="extrapolate") interpolated = f(macro_vals) # Enforce strict monotonicity for i in range(1, len(interpolated)): if increasing: if interpolated[i] <= interpolated[i-1]: interpolated[i] = interpolated[i-1] + 1e-6 # Tiny increment to keep it strict else: if interpolated[i] >= interpolated[i-1]: interpolated[i] = interpolated[i-1] - 1e-6 df[f"{micro_col}_monotonic"] = interpolated return df # Apply strict increasing constraint to col1 df = interpolate_monotonic(df, "col1", "col_macro", increasing=True)
To ensure the interpolation matches macro dynamics:
- Check correlation between filled micro data and macro data:
print(df["col1_macro_interp"].corr(df["col_macro"])) - Plot the series to visually verify alignment:
import matplotlib.pyplot as plt plt.figure(figsize=(10, 6)) plt.plot(df["col_macro"], label="Macro Data", linestyle="--") plt.plot(df["col1_macro_interp"], label="Interpolated Col1") plt.plot(df["col2_const_constant"], label="Constrained Col2") plt.legend() plt.show()
内容的提问来源于stack exchange,提问作者Babak Fi Foo

