Python使用threadid并行计算时出现结果缺失问题如何解决
问题原因
- 核心触发原因:线程创建数量绑定了TensorFlow识别到的逻辑GPU数量,而非参数列表的长度。你遇到的4参数仅输出3结果的情况,本质是当前环境可被TensorFlow识别到的逻辑GPU只有3个,
len(gpu_list)返回值为3,因此循环仅启动了3个线程,仅处理了下标0、1、2对应的三个参数,下标3的第四个参数完全没有被执行。当参数为3个时刚好和GPU数量匹配,因此输出正常。 - 潜在异常问题:代码存在函数名拼写不一致问题,你定义的函数名为
runsoncores(带字母s),但在parallel_run中调用的是runoncores(不带字母s),如果实际运行的代码也存在该拼写错误,会直接抛出未定义函数的异常,也会导致线程异常退出无结果输出。
解决方案
你可以通过以下修改实现参数数量和输出结果数量一致:
- 先修正代码中的拼写错误,保证函数定义和调用的名称统一。
- 将线程数量改为和参数列表长度绑定,同时添加GPU资源复用逻辑,参数数量多于GPU数量时自动轮询分配GPU:
import numpy as np import tensorflow as tf import threading import time parameter_values = [0.2,0.3,0.4,0.5] # 统一函数名 def runoncores(threadid): np.random.seed(threadid) tf.random.set_seed(threadid) parameter_for_sim = parameter_values[threadid] # **runs simulation here with the parameter value** # samples变量为你自行生成的模拟结果,这里仅做占位示例 samples = None filename = 'results_' + str(parameter_for_sim) + '.npy' np.save(filename,samples) def parallel_run(threadid, gpu): with tf.name_scope(gpu): with tf.device(gpu): runoncores(threadid) return gpu_list = tf.config.experimental.list_logical_devices('GPU') # 线程数改为和参数列表长度绑定 num_threads = len(parameter_values) print(num_threads) threads = list() start = time.time() for index in range(num_threads): # 轮询分配GPU,参数数量多于GPU时自动复用GPU资源 assigned_gpu = gpu_list[index % len(gpu_list)].name x = threading.Thread(target=parallel_run, args=(index, assigned_gpu)) threads.append(x) x.start() for index, thread in enumerate(threads): thread.join() end = time.time() print('Threaded time taken: ', end-start)
- 如果运行中出现GPU显存不足的报错,可以添加TensorFlow显存动态分配配置,或者对参数分批执行,每批次执行的数量不超过GPU的数量。
内容的提问来源于stack exchange,提问作者Lizardinablizzard
相关产品推荐
相关产品推荐

