You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

concurrent.futures线程完成状态无法显示的问题求助

concurrent.futures线程完成状态无法显示的问题求助

各位大佬好!最近用concurrent.futures做多线程任务的时候卡壳了——线程的完成状态总是没法正常显示,进度输出要么混乱要么报错,折腾半天没搞定,来求助大家🥺

我写的代码片段如下:

import concurrent.futures
import random
import pdb

# Analysis of text packet
def Threads1(curr_section, index1):
    words = open('test.txt', 'r', encoding='utf-8', errors='ignore').read().replace('"', '').split()
    
    longest_recorded = []
    for ii1 in words:
        test1 = random.randint(1, 1000)
        if test1 > 900: break
        else: longest_recorded.append(ii1)
    
    perc = (index1 / max1) * 100
    print('In:  ' + str([index1, str(int(perc))+'%']))
    return [index1, longest_recorded]

目前遇到的问题:

  • 运行时经常报NameError,提示max1未定义,进度计算完全没法正常工作;
  • 多个线程同时print的时候,输出内容互相穿插,根本分不清哪个线程完成了;
  • 没法实时跟踪每个线程的完成状态,不知道哪些线程已经跑完,哪些还在运行。

想请教下各位大佬:

  1. 我的代码里哪些地方写得有问题?
  2. 有没有靠谱的方法可以清晰显示每个线程的完成状态和整体进度?

可能的解决思路和修正示例

我查了些资料,整理了几个改进方向:

  1. 提前读取文件,避免重复IO:每个线程都重新打开读取文件太浪费资源,还可能引发IO竞争,建议先把文件内容读到全局变量里;
  2. 明确线程总数或用线程安全计数器:max1要提前定义好(比如线程池的大小),或者用threading.Lock维护全局计数器,确保进度计算准确;
  3. 用as_completed跟踪线程完成状态:通过concurrent.futures.as_completed()逐个获取完成的线程,有序输出完成信息,不会混乱;
  4. 避免随机中断干扰状态判断:如果不是业务必须,尽量不要用随机break,不然线程提前退出会让状态跟踪变得混乱。

修正后的示例代码大概是这样的:

import concurrent.futures
import random
import threading

# 提前读取文件内容,避免每个线程重复读取
with open('test.txt', 'r', encoding='utf-8', errors='ignore') as f:
    global_words = f.read().replace('"', '').split()

# 线程安全的计数器和锁
completed_count = 0
count_lock = threading.Lock()
total_threads = 5  # 假设我们开5个线程

def Threads1(curr_section, index1):
    global completed_count, count_lock, total_threads
    longest_recorded = []
    for ii1 in global_words:
        test1 = random.randint(1, 1000)
        if test1 > 900: 
            break
        longest_recorded.append(ii1)
    
    # 线程安全更新完成计数
    with count_lock:
        completed_count += 1
        perc = (completed_count / total_threads) * 100
        print(f"线程 {index1} 已完成!当前整体进度:{int(perc)}%")
    
    return [index1, longest_recorded]

# 使用线程池并跟踪完成状态
if __name__ == "__main__":
    with concurrent.futures.ThreadPoolExecutor(max_workers=total_threads) as executor:
        # 提交所有线程任务
        futures = {executor.submit(Threads1, i, i): i for i in range(total_threads)}
        
        # 逐个获取完成的线程结果
        for future in concurrent.futures.as_completed(futures):
            thread_idx = futures[future]
            try:
                result = future.result()
                print(f"线程 {thread_idx} 返回结果:{result[:10]}...")  # 只打印部分结果避免过长
            except Exception as exc:
                print(f"线程 {thread_idx} 执行出错:{exc}")

这样修改后,应该就能清晰看到每个线程的完成状态和整体进度,也不会出现输出混乱的情况了。


备注:内容来源于stack exchange,提问作者Rhys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 17:07:58