You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

强化学习中折扣奖励和的高效计算:如何无循环实现?

无循环实现强化学习折扣奖励和

问题描述

给定奖励列表 [1, 2, 3] 和折扣值 0.9,需完成:

  1. 生成折扣数组:[1, 0.9, 0.9²]
  2. 计算折扣奖励和数组,每个元素为 rewards[i:] 与 discounts[:n-i] 的逐元素乘积之和。

现有带循环的实现:

def get_discounted_rewards(rewards, discount):
    n = len(rewards)
    sum_of_discounted_rewards = []
    rewards = np.array(rewards)
    discounts = discount ** np.arange(n)
    for i in range(n):
        sum_of_discounted_rewards.append(np.sum(rewards[i:] * discounts[:n-i]))
    return sum_of_discounted_rewards

求无显式/隐式循环的实现方式。


实现方案

方案1:下三角矩阵广播计算

通过构造矩阵实现全向量化运算,完全避开循环:

import numpy as np

def get_discounted_rewards_no_loop(rewards, discount):
    rewards = np.array(rewards)
    n = len(rewards)
    discounts = discount ** np.arange(n)
    
    # 构造奖励下三角矩阵:每行i仅保留rewards[i:],其余位置补0
    reward_mat = np.tril(np.tile(rewards, (n, 1)))
    # 构造折扣匹配矩阵:每行i仅保留前n-i个折扣值
    discount_mat = np.triu(discounts).T
    
    # 逐元素乘积后按行求和,转换为列表返回
    return np.sum(reward_mat * discount_mat, axis=1).tolist()

方案2:逆向累积求和(高效O(n)实现)

从奖励数组末尾开始逆向计算,利用numpy的向量化累积操作,计算复杂度远低于矩阵方法:

import numpy as np

def get_discounted_rewards_no_loop(rewards, discount):
    rewards = np.array(rewards, dtype=np.float64)
    n = len(rewards)
    
    # 逆序奖励数组,计算带折扣因子的累积和
    reversed_rewards = rewards[::-1]
    discount_pows = discount ** np.arange(n)
    cumulative = np.cumsum(reversed_rewards * discount_pows)
    
    # 修正折扣因子并逆序,得到最终结果
    return (cumulative[::-1] / discount_pows).tolist()

结果验证

输入rewards=[1,2,3]、discount=0.9时,两种方案均输出:
[5.23, 4.7, 3],与原循环代码结果完全一致。


内容的提问来源于stack exchange,提问作者max_max_mir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 04:57:17