You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

场函数元模型实现遇异常:样本数与多项式阶数相关误差反常

场函数元模型实现异常问题排查

基于系统输入输出文件实现场函数元模型,但出现以下异常结果:

输入输出文件说明

  • 输入文件含5列(对应系统5个参数),共nsamples行(拉丁超立方样本数);
  • 输出文件含nsamples列,共ntimesteps行。

异常问题

  1. 误差百分比反常:样本数更多时,误差反而上升,不符合预期;
  2. 参数维度影响显著:3参数系统中,多项式阶数为6时误差0.12;但5参数系统中,最优误差为阶数3时的1.2,阶数6时误差极大。

实现代码

##########################################################################################

nsamples  = 50 # The number of samples
ntimesteps = 111 # the number of timestep for each simulation in the model 
num_param = 5 # number of parameters of the model 

inputs = pd.read_csv(f"samples_{nsamples}.csv", usecols=range(1, num_param+1), skipinitialspace=True)
outputs = pd.read_csv(f'output_{nsamples}.csv', skiprows=1, skipinitialspace=True, header=None )
outputs = outputs.transpose()

mesh = ot.IntervalMesher([ntimesteps - 1]).build(ot.Interval(0, ntimesteps - 1)) # construction a one-dimensional mesh

sample_X = ot.Sample(np.array(inputs.iloc[:,0:num_param]))
X = np.array(outputs.iloc[0,:])
X = X.reshape((ntimesteps,1))

Field = ot.Field(mesh,X) # field with the just first row
sample_Y = ot.ProcessSample(1,Field) # It stores a list of samples in which each value of the mesh is associated to an output value

for k in range(1,nsamples):
    Xi = np.array(outputs.iloc[k,:]).reshape(ntimesteps,1)
    New_Field = ot.Field(mesh,Xi)
    sample_Y.add(New_Field)


threshold = 1e-7
algo_Y = ot.KarhunenLoeveSVDAlgorithm(sample_Y, threshold)
algo_Y.run()
result_Y = algo_Y.getResult()
postProcessing = ot.KarhunenLoeveLifting(result_Y)
outputSampleChaos = result_Y.project(sample_Y)


degree = 6
dimension_xi_X = num_param
dimension_xi_Y = ntimesteps
enumerateFunction = ot.HyperbolicAnisotropicEnumerateFunction(dimension_xi_X, 0.8)
basis = ot.OrthogonalProductPolynomialFactory(
    [ot.StandardDistributionPolynomialFactory(ot.HistogramFactory().build(sample_X[:,i])) for i in range(dimension_xi_X)], enumerateFunction)
basisSize = enumerateFunction.getStrataCumulatedCardinal(degree)
adaptive = ot.FixedStrategy(basis, basisSize)
projection = ot.LeastSquaresStrategy(
    ot.LeastSquaresMetaModelSelectionFactory(ot.LARS(), ot.CorrectedLeaveOneOut()))
ot.ResourceMap.SetAsScalar("LeastSquaresMetaModelSelection-ErrorThreshold", 1.0e-7)
algo_chaos = ot.FunctionalChaosAlgorithm(sample_X,
outputSampleChaos,basis.getMeasure(), adaptive, projection)
algo_chaos.run()
metaModel1 = ot.PointToFieldConnection(postProcessing, algo_chaos.getResult().getMetaModel())

def metaModel(theta):
    out=metaModel1(theta)
    return(((out)[0:ntimesteps]))## Change the number of outputs
 
error = rmse_error() # this is a function that calculate the difference between the model's outputs and the metamodel in percentage
print(error)
##########################################################################################

问题分析与解决建议

针对异常1:样本数增加反而误差上升

  • 核心原因:
    1. 样本质量缺陷:拉丁超立方采样可能混入重复点或未覆盖参数空间关键区域,新增样本未补充有效信息反而引入噪声;
    2. 过拟合风险:样本数增加但模型复杂度未匹配调整时,易出现过拟合,尤其高维场景下更明显;
    3. 误差计算逻辑错误:若rmse_error()用训练集计算误差,样本数增加会让模型拟合训练数据更充分,但泛化误差可能上升,误将训练误差当作泛化误差会导致误解。
  • 解决建议:
    • 校验采样质量:检查新增样本的参数分布,确保均匀覆盖参数空间,剔除冗余或极端噪声点;
    • 拆分数据集:将样本划分为训练集(70%-80%)和测试集(20%-30%),仅用训练集训练模型,测试集计算泛化误差;
    • 匹配模型复杂度:随样本数增加,适当提升模型复杂度,同时用交叉验证监控泛化误差。

针对异常2:5参数系统高多项式阶数误差激增

  • 核心原因:
    1. 维度诅咒:参数维度从3升至5时,多项式基函数数量随阶数指数增长,阶数6时的基函数规模远大于3参数场景,直接导致过拟合;
    2. 样本量不足:5参数+阶数6的模型需要大量样本支撑训练,当前50个样本无法满足系数稳定估计的需求;
    3. 枚举函数设置不合理:HyperbolicAnisotropicEnumerateFunction的0.8参数未有效限制高交互项数量,导致基函数冗余度高。
  • 解决建议:
    • 降低多项式阶数:针对5参数系统,优先选择阶数3-4,或替换FixedStrategy为AdaptiveStrategy,自动筛选重要基函数;
    • 扩充样本量:高维高复杂度模型需更多样本,建议将样本量提升至200以上,确保样本数至少为基函数数量的5-10倍;
    • 调整枚举函数:降低HyperbolicAnisotropicEnumerateFunction的第二个参数(如设为0.5),减少高交互项生成,控制基函数规模;
    • 强化正则化:在LARS算法基础上增加更严格的正则约束(如Lasso),限制系数大小,缓解过拟合。

内容的提问来源于stack exchange,提问作者hndphm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 17:03:19