理论高效的张量操作Routine运行慢于未优化版,求优化Routine2
4维张量自收缩后截断膨胀索引的优化问题
本次目标是操作4维张量,完成自收缩后将膨胀索引截断回原长度。现有两种实现方案:
- 方案1:内存复杂度为
dimension^5,需生成形状为(dimension,dimension,dimension,dimension,dimension)的tensorB2和tensorC,当dimension=60时耗时约66秒。 - 方案2:计算结果与方案1完全一致,通过拆分计算避免生成超过
(dimension,dimension,dimension,dimension)的张量,内存复杂度为dimension^4,理论性能更优,但dimension=60时耗时约30分钟。
已确认两者计算结果完全相同,怀疑是numpy ndarray使用不当导致内存冗余或无效操作,寻求方案2的优化思路。
导入模块与辅助函数
import numpy as np from ncon import ncon def truncate_matrix(matrix, index, dim): shape = list(matrix.shape) shape[index] = dim newMatrix = matrix[0:shape[0], 0:shape[1]] return newMatrix
方案1(内存复杂度dimension^5)
dimension = 60 myTensor = np.random.rand(dimension, dimension, dimension, dimension) tensorA = ncon([myTensor, myTensor], [[-1,1,2,-3], [-2,1,2,-4]]) tensorB = ncon([myTensor, myTensor], [[-1,1,-3,2], [-2,1,-4,2]]) tensorQ = ncon([tensorA, tensorB], [[-1,-3,1,2], [-2,-4,1,2]]) del tensorA, tensorB # 求解Utr Qshape = tensorQ.shape tensorQ = tensorQ.reshape((Qshape[0]*Qshape[1], Qshape[2]*Qshape[3])) Ul, deltaL, _ = np.linalg.svd(tensorQ) Utr = np.array(truncate_matrix(Ul, 1, dimension)) Utr = Utr.reshape((Qshape[0], Qshape[1], dimension)) del tensorQ, Ul, Qshape # 重构张量T tensorB2 = ncon([Utr, myTensor], [[1,-5,-1], [1,-4,-2,-3]]) tensorC = ncon([tensorB2, myTensor], [[-1,-2,1,-4,2], [2,-5,1,-3]]) del tensorB2 newTensor1 = ncon([Utr, tensorC], [[1,2,-2], [-1,-3,-4,1,2]]) del tensorC
方案2(内存复杂度dimension^4)
dimension = 60 myTensor = np.random.rand(dimension, dimension, dimension, dimension) tensorA = ncon([myTensor, myTensor], [[-1,1,2,-3], [-2,1,2,-4]]) tensorB = ncon([myTensor, myTensor], [[-1,1,-3,2], [-2,1,-4,2]]) tensorQ = ncon([tensorA, tensorB], [[-1,-3,1,2], [-2,-4,1,2]]) # 求解Utr Qshape = tensorQ.shape tensorQ = tensorQ.reshape((Qshape[0]*Qshape[1], Qshape[2]*Qshape[3])) Ul, deltaL, _ = np.linalg.svd(tensorQ) Utr = np.array(truncate_matrix(Ul, 1, dimension)) Utr = Utr.reshape((Qshape[0], Qshape[1], dimension)) courrentShape = myTensor.shape newTensor2 = np.zeros((dimension, dimension, courrentShape[2], courrentShape[3])) # 重构张量T for xf in range(0, newTensor2.shape[0]): for yu in range(0, dimension): # 切片当前张量 tensorSlice = myTensor[:,:,yu,:] # T_iy1y'1 # 切片截断矩阵 matrixSlice = Utr[:,:,xf] # Uy1y2 tensorB2 = ncon([matrixSlice, tensorSlice], [[1,-3], [1,-2,-1]]) for yb in range(0, newTensor2.shape[3]): tensorSlice = myTensor[:,:,:,yb] tensorC = ncon([tensorB2, tensorSlice], [[1,-1,2], [2,-2,1]]) newTensor2[xf,:,yu,yb] = ncon([Utr, tensorC], [[1,2,-1], [1,2]])
方案2的优化思路
1. 消除嵌套循环,改用向量化操作
方案2的三重嵌套循环是性能瓶颈核心——Python循环本身开销极大,dimension=60时总循环次数达216000次,每次循环调用ncon又会额外增加张量操作开销。直接将重构逻辑转化为numpy批量张量运算:
# 替代三重循环的向量化实现 # 第一步:批量计算 tensorB2 tensorB2_batch = np.einsum('abc,acd->abd', Utr, myTensor) # 第二步:批量计算 tensorC tensorC_batch = np.einsum('abcd,ace->abde', tensorB2_batch, myTensor) # 第三步:生成最终张量 newTensor2 = np.einsum('abc,abde->cde', Utr, tensorC_batch) # 及时释放临时大张量 del tensorB2_batch, tensorC_batch
2. 减少ncon频繁调用的开销
ncon每次调用都会进行索引解析和维度调整,循环内反复调用会累积不必要的损耗。改用numpy原生的einsum或tensordot,直接指定收缩模式,减少中间环节的性能浪费。
3. 验证计算一致性
优化后可通过以下代码确认结果与原方案一致:
print(np.allclose(newTensor1, newTensor2)) # 输出应为True
内容的提问来源于stack exchange,提问作者Indiano
相关产品推荐
相关产品推荐

