You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何调用concurrent.futures传入列表与非列表双参数

问题根因

代码运行异常的核心原因是对ThreadPoolExecutor.map()的参数规则使用错误:

  • executor.map(目标函数, 入参可迭代对象1, 入参可迭代对象2, ...)的执行逻辑是按位置,并行从每个传入的可迭代对象中取元素作为目标函数的入参,要求除目标函数外,其余传入的参数都必须是可迭代类型,且任务分发会在最短的可迭代对象遍历完成时终止。
  • 你传入的cursor是pymssql游标实例,不属于可迭代类型,直接传入会触发类型错误;即便手动将其包装为可迭代对象,也会因为长度和ssns列表不匹配,导致任务执行数量不符合预期。
  • 额外隐藏风险:pymssql的连接、游标对象非线程安全,跨线程共享同一个游标执行数据库写入,会触发数据错乱、连接断开等不可预期的问题。
修正方案

方案1:保留map调用写法,固定参数用itertools.repeat包装

如果你的get_pdf_multi_thread函数仅读取传入的cursor配置,不会直接调用cursor执行SQL操作,可以通过itertools.repeat将固定参数包装为无限重复的可迭代对象,适配map的参数要求:

import time
import concurrent.futures
import itertools
import pymssql
# 提前导入业务依赖:f模块、db配置、create_folders函数

def main(ssns):
    t1 = time.perf_counter()
    create_folders()

    conn = pymssql.connect(db.server, db.user, db.password, db.db)
    cursor = conn.cursor(as_dict=True)

    with concurrent.futures.ThreadPoolExecutor() as executor:
        # itertools.repeat会生成无限重复cursor的可迭代对象,自动对齐ssns的遍历长度
        results = executor.map(f.get_pdf_multi_thread, ssns, itertools.repeat(cursor))
        # 需要获取所有任务返回结果时,可转换为list触发结果遍历
        # results = list(results)

    conn.commit()
    cursor.close()
    conn.close()
    print(f"任务总耗时: {time.perf_counter() - t1:.2f}s")

方案2:子线程独立创建数据库连接(推荐,规避线程安全问题)

如果子线程内需要执行数据库读写操作,禁止跨线程共享连接、游标对象,需要在每个子线程内部独立创建数据库连接,从根源上避免竞态问题:

import time
import concurrent.futures
import pymssql
# 提前导入业务依赖:f模块、db配置、create_folders函数

def thread_worker(ssn):
    # 每个工作线程独立创建连接、游标,线程退出前主动释放资源
    thread_conn = pymssql.connect(db.server, db.user, db.password, db.db)
    thread_cursor = thread_conn.cursor(as_dict=True)
    try:
        task_result = f.get_pdf_multi_thread(ssn, thread_cursor)
        thread_conn.commit()
        return task_result
    except Exception as e:
        thread_conn.rollback()
        raise e
    finally:
        thread_cursor.close()
        thread_conn.close()

def main(ssns):
    t1 = time.perf_counter()
    create_folders()

    with concurrent.futures.ThreadPoolExecutor() as executor:
        results = list(executor.map(thread_worker, ssns))

    print(f"任务总耗时: {time.perf_counter() - t1:.2f}s")
注意事项
  • 绝大多数Python生态的数据库驱动(pymssql、pymysql、sqlite3等)的连接、游标对象都未做线程安全适配,多线程场景下禁止跨线程共享这类资源。
  • 如果子线程逻辑不需要操作数据库,不要将连接、游标作为参数传入,减少不必要的资源传递。

内容的提问来源于stack exchange,提问作者Andoni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 19:57:16