You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

理论高效的张量操作Routine运行慢于未优化版,求优化Routine2

4维张量自收缩后截断膨胀索引的优化问题

本次目标是操作4维张量,完成自收缩后将膨胀索引截断回原长度。现有两种实现方案:

  • 方案1:内存复杂度为dimension^5,需生成形状为(dimension,dimension,dimension,dimension,dimension)的tensorB2和tensorC,当dimension=60时耗时约66秒。
  • 方案2:计算结果与方案1完全一致,通过拆分计算避免生成超过(dimension,dimension,dimension,dimension)的张量,内存复杂度为dimension^4,理论性能更优,但dimension=60时耗时约30分钟。

已确认两者计算结果完全相同,怀疑是numpy ndarray使用不当导致内存冗余或无效操作,寻求方案2的优化思路。


导入模块与辅助函数

import numpy as np 
from ncon import ncon

def truncate_matrix(matrix, index, dim):
    shape = list(matrix.shape)
    shape[index] = dim
    newMatrix = matrix[0:shape[0], 0:shape[1]]
    return newMatrix

方案1(内存复杂度dimension^5)

dimension = 60 
myTensor = np.random.rand(dimension, dimension, dimension, dimension)
tensorA = ncon([myTensor, myTensor], [[-1,1,2,-3], [-2,1,2,-4]])
tensorB = ncon([myTensor, myTensor], [[-1,1,-3,2], [-2,1,-4,2]])
tensorQ = ncon([tensorA, tensorB], [[-1,-3,1,2], [-2,-4,1,2]])
del tensorA, tensorB

# 求解Utr
Qshape = tensorQ.shape
tensorQ = tensorQ.reshape((Qshape[0]*Qshape[1], Qshape[2]*Qshape[3]))
Ul, deltaL, _ = np.linalg.svd(tensorQ)
Utr = np.array(truncate_matrix(Ul, 1, dimension))
Utr = Utr.reshape((Qshape[0], Qshape[1], dimension))
del tensorQ, Ul, Qshape

# 重构张量T
tensorB2 = ncon([Utr, myTensor], [[1,-5,-1], [1,-4,-2,-3]])
tensorC = ncon([tensorB2, myTensor], [[-1,-2,1,-4,2], [2,-5,1,-3]])
del tensorB2
newTensor1 = ncon([Utr, tensorC], [[1,2,-2], [-1,-3,-4,1,2]])          
del tensorC

方案2(内存复杂度dimension^4)

dimension = 60 
myTensor = np.random.rand(dimension, dimension, dimension, dimension)
tensorA = ncon([myTensor, myTensor], [[-1,1,2,-3], [-2,1,2,-4]])
tensorB = ncon([myTensor, myTensor], [[-1,1,-3,2], [-2,1,-4,2]])
tensorQ = ncon([tensorA, tensorB], [[-1,-3,1,2], [-2,-4,1,2]])

# 求解Utr
Qshape = tensorQ.shape
tensorQ = tensorQ.reshape((Qshape[0]*Qshape[1], Qshape[2]*Qshape[3]))
Ul, deltaL, _ = np.linalg.svd(tensorQ)
Utr = np.array(truncate_matrix(Ul, 1, dimension))
Utr = Utr.reshape((Qshape[0], Qshape[1], dimension))
courrentShape = myTensor.shape
newTensor2 = np.zeros((dimension, dimension, courrentShape[2], courrentShape[3]))

# 重构张量T
for xf in range(0, newTensor2.shape[0]):
    for yu in range(0, dimension):
        # 切片当前张量
        tensorSlice = myTensor[:,:,yu,:] # T_iy1y'1
        # 切片截断矩阵
        matrixSlice = Utr[:,:,xf] # Uy1y2
        tensorB2 = ncon([matrixSlice, tensorSlice], [[1,-3], [1,-2,-1]]) 
        for yb in range(0, newTensor2.shape[3]):
            tensorSlice = myTensor[:,:,:,yb]
            tensorC = ncon([tensorB2, tensorSlice], [[1,-1,2], [2,-2,1]])
            newTensor2[xf,:,yu,yb] = ncon([Utr, tensorC], [[1,2,-1], [1,2]])

方案2的优化思路

1. 消除嵌套循环,改用向量化操作

方案2的三重嵌套循环是性能瓶颈核心——Python循环本身开销极大,dimension=60时总循环次数达216000次,每次循环调用ncon又会额外增加张量操作开销。直接将重构逻辑转化为numpy批量张量运算:

# 替代三重循环的向量化实现
# 第一步:批量计算 tensorB2
tensorB2_batch = np.einsum('abc,acd->abd', Utr, myTensor)

# 第二步:批量计算 tensorC
tensorC_batch = np.einsum('abcd,ace->abde', tensorB2_batch, myTensor)

# 第三步:生成最终张量
newTensor2 = np.einsum('abc,abde->cde', Utr, tensorC_batch)

# 及时释放临时大张量
del tensorB2_batch, tensorC_batch

2. 减少ncon频繁调用的开销

ncon每次调用都会进行索引解析和维度调整,循环内反复调用会累积不必要的损耗。改用numpy原生的einsum或tensordot,直接指定收缩模式,减少中间环节的性能浪费。

3. 验证计算一致性

优化后可通过以下代码确认结果与原方案一致:

print(np.allclose(newTensor1, newTensor2))  # 输出应为True

内容的提问来源于stack exchange,提问作者Indiano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 21:01:36