基于正弦波的RL交易代理无合理行为,求问题排查建议
基于正弦波价格序列的RL交易代理训练困境
正在构建一个基于合成正弦波价格序列的强化学习(RL)交易代理,这是能想到的最简单测试数据集,用于验证代理是否能学会低买高卖。但经过数周实验,模型仍未展现出有意义的行为。以下是实验设置与尝试:
🧪 环境设置
- 基于gymnasium.Env构建简单环境,生成正弦波价格序列,可添加噪声
- 动作空间包含3种离散动作:0 = HOLD(持有)、1 = ENTER(开仓)、2 = EXIT(平仓)
- 奖励基于盈亏(PnL)百分比,同时尝试过针对过度交易或无效动作设置不同力度的惩罚
- 观测特征包括:
norm_dist:价格在正弦波极值间归一化后的值proximity_to_extremes:价格与正弦波极值的接近程度log_return:时间步间的对数收益率hold_duration:持仓持续步数action_hint:甚至添加了直接的动作提示
🧠 尝试的模型
Dueling DQN(基于Stable-Baselines3)- 采用
MlpPolicy的PPO - 采用
MlpLstmPolicy的RecurrentPPO(基于SB3-contrib) - 尝试过大量超参数,迭代次数最高达1,500,000
实验结果
- 模型确实学会了开仓和平仓,但时机选择毫无逻辑
- 无论价格涨跌都会开仓,平仓也不考虑盈亏情况
- 有时交易频率会变化,但交易逻辑仍无改进
- 尝试过单步输入观测和滚动窗口输入,均未提升性能
额外调试步骤
- 在正弦波上可视化交易,仍为随机开平仓
- 打印每一步的动作、奖励和盈亏数据
- 记录TensorBoard指标,无明显收敛迹象
- 验证奖励值非零,且与优质交易正确关联
恳请各位提供反馈与见解,以下是相关代码片段:
环境(简化版)
class SineTradeEnv(gym.Env): def __init__(self, prices, norm_dist, proximity, log_return): self.prices = prices self.norm_dist_array = norm_dist self.proximity_array = proximity self.log_return_array = log_return self.max_t = len(prices) self.action_space = spaces.Discrete(3) # HOLD, OPEN, CLOSE self.observation_space = spaces.Box(low=-np.inf, high=np.inf, shape=(5,), dtype=np.float32) def reset(self): self.t = 0 self.position = None self.entry_step = None return self._get_obs(), {} def step(self, action): self.t += 1 done = self.t >= self.max_t - 1 current_price = self.prices[self.t] # PnL-based reward if self.position is not None: pnl_pct = (current_price - self.position) / self.position else: pnl_pct = 0 # Reward logic reward = 0 if self.position is None and action == 1: self.position = current_price self.entry_step = self.t elif self.position is not None and action == 2: reward = pnl_pct * 3 if pnl_pct > 0 else pnl_pct self.position = None self.entry_step = None elif self.position is not None and action == 0: reward = pnl_pct * 1 return self._get_obs(), reward, done, False, {} def _get_obs(self): return np.array([ float(self.position is not None), self.norm_dist_array[self.t], self.proximity_array[self.t], self.log_return_array[self.t], float(self.t - self.entry_step) if self.entry_step is not None else 0.0 ], dtype=np.float32)
模型训练(Stable-Baselines3 PPO)
model = PPO( policy="MlpPolicy", env=env, n_steps=1024, batch_size=512, learning_rate=2.5e-4, ent_coef=0.01, verbose=1, tensorboard_log="./tensorboard_logs/" ) model.learn(total_timesteps=1_500_000)
数据生成(合成正弦波)
def generate_price_series(max_t=500): x = np.linspace(0, 6 * np.pi, max_t) sine = np.sin(x) prices = 1000 + 200 * sine log_return = np.append([0], np.diff(np.log(prices))) norm_dist = (prices - prices.min()) / (prices.max() - prices.min()) * 2 - 1 proximity = np.minimum( (prices - prices.min()) / (prices.max() - prices.min()), (prices.max() - prices) / (prices.max() - prices.min()) ) return prices, norm_dist, proximity, log_return
推理与调试可视化
obs, _ = env.reset() for t in range(env.max_t): action, _ = model.predict(obs, deterministic=True) obs, reward, done, _, _ = env.step(int(action)) print(f"t={t} | action={action} | reward={reward:.4f}") if done: break
内容的提问来源于stack exchange,提问作者Oleg Bizin
相关产品推荐
相关产品推荐

