You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于正弦波的RL交易代理无合理行为,求问题排查建议

基于正弦波价格序列的RL交易代理训练困境

正在构建一个基于合成正弦波价格序列的强化学习(RL)交易代理,这是能想到的最简单测试数据集,用于验证代理是否能学会低买高卖。但经过数周实验,模型仍未展现出有意义的行为。以下是实验设置与尝试:

🧪 环境设置

  • 基于gymnasium.Env构建简单环境,生成正弦波价格序列,可添加噪声
  • 动作空间包含3种离散动作:0 = HOLD(持有)、1 = ENTER(开仓)、2 = EXIT(平仓)
  • 奖励基于盈亏(PnL)百分比,同时尝试过针对过度交易或无效动作设置不同力度的惩罚
  • 观测特征包括:
    • norm_dist:价格在正弦波极值间归一化后的值
    • proximity_to_extremes:价格与正弦波极值的接近程度
    • log_return:时间步间的对数收益率
    • hold_duration:持仓持续步数
    • action_hint:甚至添加了直接的动作提示

🧠 尝试的模型

  • Dueling DQN(基于Stable-Baselines3)
  • 采用MlpPolicy的PPO
  • 采用MlpLstmPolicy的RecurrentPPO(基于SB3-contrib)
  • 尝试过大量超参数,迭代次数最高达1,500,000

实验结果

  • 模型确实学会了开仓和平仓,但时机选择毫无逻辑
  • 无论价格涨跌都会开仓,平仓也不考虑盈亏情况
  • 有时交易频率会变化,但交易逻辑仍无改进
  • 尝试过单步输入观测和滚动窗口输入,均未提升性能

额外调试步骤

  • 在正弦波上可视化交易,仍为随机开平仓
  • 打印每一步的动作、奖励和盈亏数据
  • 记录TensorBoard指标,无明显收敛迹象
  • 验证奖励值非零,且与优质交易正确关联

恳请各位提供反馈与见解,以下是相关代码片段:

环境(简化版)

class SineTradeEnv(gym.Env):
    def __init__(self, prices, norm_dist, proximity, log_return):
        self.prices = prices
        self.norm_dist_array = norm_dist
        self.proximity_array = proximity
        self.log_return_array = log_return
        self.max_t = len(prices)
        self.action_space = spaces.Discrete(3)  # HOLD, OPEN, CLOSE
        self.observation_space = spaces.Box(low=-np.inf, high=np.inf, shape=(5,), dtype=np.float32)

    def reset(self):
        self.t = 0
        self.position = None
        self.entry_step = None
        return self._get_obs(), {}

    def step(self, action):
        self.t += 1
        done = self.t >= self.max_t - 1
        current_price = self.prices[self.t]

        # PnL-based reward
        if self.position is not None:
            pnl_pct = (current_price - self.position) / self.position
        else:
            pnl_pct = 0

        # Reward logic
        reward = 0
        if self.position is None and action == 1:
            self.position = current_price
            self.entry_step = self.t
        elif self.position is not None and action == 2:
            reward = pnl_pct * 3 if pnl_pct > 0 else pnl_pct
            self.position = None
            self.entry_step = None
        elif self.position is not None and action == 0:
            reward = pnl_pct * 1

        return self._get_obs(), reward, done, False, {}

    def _get_obs(self):
        return np.array([
            float(self.position is not None),
            self.norm_dist_array[self.t],
            self.proximity_array[self.t],
            self.log_return_array[self.t],
            float(self.t - self.entry_step) if self.entry_step is not None else 0.0
        ], dtype=np.float32)

模型训练(Stable-Baselines3 PPO)

model = PPO(
    policy="MlpPolicy",
    env=env,
    n_steps=1024,
    batch_size=512,
    learning_rate=2.5e-4,
    ent_coef=0.01,
    verbose=1,
    tensorboard_log="./tensorboard_logs/"
)

model.learn(total_timesteps=1_500_000)

数据生成(合成正弦波)

def generate_price_series(max_t=500):
    x = np.linspace(0, 6 * np.pi, max_t)
    sine = np.sin(x)
    prices = 1000 + 200 * sine
    log_return = np.append([0], np.diff(np.log(prices)))
    norm_dist = (prices - prices.min()) / (prices.max() - prices.min()) * 2 - 1
    proximity = np.minimum(
        (prices - prices.min()) / (prices.max() - prices.min()),
        (prices.max() - prices) / (prices.max() - prices.min())
    )
    return prices, norm_dist, proximity, log_return

推理与调试可视化

obs, _ = env.reset()
for t in range(env.max_t):
    action, _ = model.predict(obs, deterministic=True)
    obs, reward, done, _, _ = env.step(int(action))
    print(f"t={t} | action={action} | reward={reward:.4f}")
    if done:
        break

内容的提问来源于stack exchange,提问作者Oleg Bizin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 17:23:20