自定义Autodiff梯度计算实现与TensorFlow结果不符求助
自动微分matmul梯度计算与TensorFlow结果不符问题排查
问题背景
我在学习自动微分(Autodiff)算法时,自行查阅论文实现了一套逻辑,但测试结果与TensorFlow输出大多不符。之后参考相关教程,仅针对矩阵乘法(matmul)操作改用TensorFlow算子实现,但结果仍和最初版本一致,无法匹配TensorFlow的输出。
实现代码
matmul梯度计算与unbroadcast方法
def gradient_matmul(node, dx, adj): # dx is needed to know which of both parents should be derived a = node.parents[0] b = node.parents[1] # the operation was node.tensor = tf.matmul(a.tensor, b.tensor) if a == dx or b == dx: # result depends on which of the parents is the derivative mm = tf.matmul(adj, tf.transpose(b.tensor)) if a == dx else \ tf.matmul(tf.transpose(a.tensor), adj) return mm else: return None def unbroadcast(adjoint, node): dim_a = len(adjoint.shape) dim_b = len(node.shape) if dim_a > dim_b: sum = tuple(range(dim_a - dim_b)) res = tf.math.reduce_sum(adjoint, axis = sum) return res return adjoint
自动微分梯度计算主算法
def gradient(y, dx): working = [y] adjoints = defaultdict(float) adjoints[y] = tf.ones(y.tensor.shape) while len(working) != 0: curr = working.pop(0) if curr == dx: return adjoints[curr] if curr.is_store: continue adj = adjoints[curr] for p in curr.parents: # for testing with matrix multiplication as only operation local_grad = gradient_matmul(curr, p, adj) adjoints[p] = unbroadcast(tf.add(adjoints[p], local_grad), p.tensor) if not p in working: working.append(p)
测试用例
x = tf.constant([[[1.0, 1.0], [2.0, 3.0]], [[4.0, 5.0], [6.0, 7.0]]]) y = tf.constant([[3.0, -7.0], [-1.0, 5.0]]) z = tf.constant([[[1, 1], [2.0, 2]], [[3, 3], [-1, -1]]]) w = tf.matmul(tf.matmul(x, y), z)
结果对比
TensorFlow计算的梯度结果
[<tf.Tensor: shape=(2, 2, 2), dtype=float32, numpy= array([[[-22., 18.], [-22., 18.]], [[ 32., -16.], [ 32., -16.]]], dtype=float32)>, <tf.Tensor: shape=(2, 2), dtype=float32, numpy= array([[66., -8.], [80., -8.]], dtype=float32)>, <tf.Tensor: shape=(2, 2, 2), dtype=float32, numpy= array([[[ 5., 5.], [ -1., -1.]], [[ 18., 18.], [-10., -10.]]], dtype=float32)>]
我的实现计算结果
[[[-5. 7.] [-5. 7.]] [[-5. 7.] [-5. 7.]]] [[33. 22.] [54. 36.]] [[[ 9. 9.] [14. 14.]] [[-5. -5.] [-6. -6.]]]
疑问
我怀疑问题出在numpy的dot与TensorFlow的matmul的差异上,但不知道如何针对TensorFlow修正梯度或unbroadcast逻辑,希望帮忙排查问题!
内容的提问来源于stack exchange,提问作者Frobeniusnorm
相关产品推荐
相关产品推荐

