动态优化程序中状态变量非负约束失效问题求助
问题分析与解决方案
你的有限时域动态优化程序中,资源非负约束失效的核心问题在于约束函数的定义方式错误,导致优化器无法正确识别并强制执行所有时刻的资源非负要求。
具体问题点
当前的resource_constraint函数返回一个标量值:
- 若所有时刻资源非负,返回
0 - 若某时刻资源变负,返回该时刻的负资源值
但Scipy的SLSQP优化器处理标量不等式约束时,仅能识别单一条件。而你的需求是所有时刻的剩余资源都必须≥0,这是40个独立的不等式约束(对应40个时间步)。当前的标量返回方式无法将所有约束传递给优化器,且梯度信息不完整,导致优化器无法正确调整控制量来满足所有时刻的约束。
修正方案
修改约束函数,使其返回一个数组,数组中每个元素对应一个时间步的剩余资源值。SLSQP会自动将数组的每个元素视为独立的不等式约束,要求所有元素≥0,从而确保每个时刻的资源都不会变为负值。
修正后的约束函数
# Non-negativity constraint for finite resource def resource_constraint(controls): finite_resource = initial_state[0] constraints = [] for t in range(horizon): finite_resource -= controls[t] constraints.append(finite_resource) return np.array(constraints)
约束调用无需修改
原有的约束定义保持不变,SLSQP会自动处理向量形式的不等式约束:
# Define constraint for finite resource finite_resource_constraint = {'type': 'ineq', 'fun': resource_constraint}
完整修正代码
import numpy as np from scipy.optimize import minimize # Problem parameters horizon = 40 # Time horizon state_dim = 2 # Dimension of the state initial_state = np.array([100.0, 0.0]) # Initial state: [finite resource, accumulated resource] # Dynamics function def dynamics_function(state, control): finite_resource = state[0] - control # Depletion of the finite resource accumulated_resource = state[1] + 0.2 * control # Accumulation of the second resource return np.array([finite_resource, accumulated_resource]) # Reward function to be maximized def reward_function(state, control): return 1 - control**2 + 2*state[1] # Example production utility, to be maximized # Non-negativity constraint for finite resource (修正版本) def resource_constraint(controls): finite_resource = initial_state[0] constraints = [] for t in range(horizon): finite_resource -= controls[t] constraints.append(finite_resource) return np.array(constraints) # Define the optimization problem def optimization_problem(controls): total_reward = 0 state = initial_state.copy() for t in range(horizon): control = controls[t] total_reward += reward_function(state, control) state = dynamics_function(state, control) return -total_reward # Maximize the total reward (minimize the negative) # Initial guess for controls initial_controls = np.zeros(horizon) # Define bounds for controls (production rate) control_bounds = (0, np.inf) # Production rate bounds with unbounded upper limit control_bounds_list = [control_bounds] * horizon # Define constraint for finite resource finite_resource_constraint = {'type': 'ineq', 'fun': resource_constraint} # Solve the optimization problem with constraints result = minimize(optimization_problem, initial_controls, method='SLSQP', bounds=control_bounds_list, constraints=finite_resource_constraint) # Extract the optimal controls (production rates) optimal_controls = result.x # Calculate the finite resource at the end of optimization final_state = initial_state.copy() for t in range(horizon): final_state = dynamics_function(final_state, optimal_controls[t]) print("Optimal Production Rates:", optimal_controls) print("Optimal Utility:", -result.fun) print("Final Finite Resource:", final_state[0])
补充说明
修正后的约束函数会为每个时间步生成一个约束条件,确保每一步的剩余资源都≥0。优化器在迭代过程中会同时满足所有这些约束,最终不会出现资源负值的情况。
内容的提问来源于stack exchange,提问作者GaL
相关产品推荐
相关产品推荐

