You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python线程堆积问题:如何避免线程执行后未被销毁?

Python多线程线程未销毁&资源泄漏问题排查

问题现象

我编写的Python多线程代码功能正常,但存在以下异常:

  • 线程返回结果后未被销毁,每次运行脚本时控制台线程编号持续递增
  • 脚本处理完成后内存占用仍持续上升,说明脚本结束后还有进程在后台运行

原代码

if __name__ == "__main__":
def run_selenium1(a, b, c, d, e):
    
    @st.cache_data(show_spinner=False)
    def get_links(i, resumeContent):
        # 业务逻辑
            for something1, something2, something3, something4, something5, something6, something7 in zip(Final_Something1, Final_Something2, Final_Something3, Final_Something4, Final_Something5, Final_Something6, Final_Something7):
                Final_Array.append((something1, something2, something3, something4, something5, something6, something7))
            driver.close()
            driver.quit()
        except:
            driver.close()
            driver.quit()


    with webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) as driver:
        try:
           # 获取links逻辑
        except:
            driver.close()
            driver.quit()

    threads = []
    for i in links:
        t = threading.Thread(target=get_links, args=(i, Content))
        t.daemon = True
        threads.append(t)
        t.start()
    for t in threads:
        t.join()
        print("Threads destroyed") # <---- 这行未打印

尝试的ThreadPoolExecutor方案

with ThreadPoolExecutor(max_workers=25) as executor:
    for i in links:
        task = executor.submit(get_links, i, resumeContent)
        task.join()
executor.shutdown()

问题根源分析

  1. 资源共享错误:get_links函数直接使用外部with块创建的WebDriver实例,多线程共享同一个driver会引发资源竞争,且with块结束时会自动销毁driver,此时线程若仍在操作driver会触发异常,导致线程无法正常结束。
  2. 缓存装饰器冲突:@st.cache_data装饰器会缓存函数执行状态与结果,导致函数关联的资源(如未释放的句柄、连接)被长期持有,无法随线程销毁而释放。
  3. 线程未正常终止:print("Threads destroyed")未打印,说明join()未执行完成,大概率是get_links内部出现未处理异常或阻塞(如driver操作卡住),导致线程一直处于运行状态。
  4. ThreadPoolExecutor使用错误:循环内逐个调用task.join()会让任务串行执行,失去多线程并行意义;且executor.shutdown()在with块外无意义(with块结束时executor会自动执行shutdown)。

修复方案

1. 线程内独立管理WebDriver资源

每个线程自行创建、销毁WebDriver实例,避免资源共享:

def get_links(i, resumeContent, result_queue):
    # 移除@st.cache_data装饰器
    driver = None
    try:
        driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
        # 执行具体业务逻辑
        # ...
        for item in zip(...):
            result_queue.put(item)
    except Exception as e:
        print(f"线程执行出错: {str(e)}")
    finally:
        if driver:
            driver.quit()  # quit会同时关闭窗口与进程,无需单独调用close

2. 线程安全的结果存储

替换全局Final_Array为queue.Queue,避免多线程数据竞争:

from queue import Queue

# 初始化线程安全队列
result_queue = Queue()

# 启动线程
threads = []
for i in links:
    t = threading.Thread(target=get_links, args=(i, Content, result_queue))
    t.start()
    threads.append(t)

# 等待所有线程结束
for t in threads:
    t.join()

# 从队列中收集结果
Final_Array = []
while not result_queue.empty():
    Final_Array.append(result_queue.get())

3. 正确使用ThreadPoolExecutor

批量提交任务,实现真正的并行执行:

from concurrent.futures import ThreadPoolExecutor

with ThreadPoolExecutor(max_workers=25) as executor:
    # 批量提交所有任务
    futures = [executor.submit(get_links, i, resumeContent, result_queue) for i in links]
    # 等待所有任务完成,可捕获异常
    for future in futures:
        try:
            future.result()
        except Exception as e:
            print(f"任务执行失败: {str(e)}")

# 收集结果
Final_Array = []
while not result_queue.empty():
    Final_Array.append(result_queue.get())

4. 规范异常捕获

避免空except,明确捕获异常类型,防止未知异常导致线程卡住:

try:
    # 业务逻辑代码
except Exception as e:
    print(f"执行异常: {str(e)}")
finally:
    # 确保资源被清理
    if driver:
        driver.quit()

内容的提问来源于stack exchange,提问作者alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 04:45:02