You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python线程间基于Queue的消息传递性能不佳(耗时数毫秒)

问题描述

我在Python中使用queue.Queue实现线程间消息传递,发现生产者入队到消费者出队的耗时超出预期:Raspberry Pi 3B+上平均耗时2ms以上,高端PC上也约为2ms。

  • 树莓派环境:Debian 11(bullseye)、内核5.15、Python 3.9.2
  • 已尝试操作:关闭其他进程、提升程序优先级,均无明显效果
  • 需求:希望优化性能(尤其针对树莓派)、排查是否存在配置问题,同时寻求合适的性能分析工具定位瓶颈

测试代码

# 测试线程间Queue消息传递耗时的程序
import time
from queue import Queue
from threading import Thread
import statistics

TIMING_TEST_LENGTH = 20     # 测试时长(秒)  

Global_Test_Queue = Queue()
# 存储耗时采样结果
Global_Timing_List = []


def test_producer_thread(out_q):
    for i in range(100):
        t = time.perf_counter()
        out_q.put_nowait(t)
        # 间隔一段时间再发送下一条消息
        time.sleep(int(TIMING_TEST_LENGTH / 100))  


def test_consumer_thread(in_q):
    while True:
        # 阻塞等待队列数据
        ticksWhenSent = in_q.get()
        
        # 记录接收时间
        tickWhenDataReceived = time.perf_counter()

        # 收到终止信号则退出线程
        if ticksWhenSent < 0:
            break

        delayMs = (tickWhenDataReceived - ticksWhenSent) * 1000
        Global_Timing_List.append(delayMs)


def main():
    try:
        print(f"开始{TIMING_TEST_LENGTH}秒的耗时测试。")

        consumerThread = Thread(target=test_consumer_thread, args=(Global_Test_Queue,))
        consumerThread.start()

        producerThread = Thread(target=test_producer_thread, args=(Global_Test_Queue,))
        producerThread.start()

        time.sleep(TIMING_TEST_LENGTH)  # 等待测试完成

        avg = statistics.mean(Global_Timing_List)
        std = statistics.stdev(Global_Timing_List)
        print(f"延迟时间(毫秒): {Global_Timing_List}")
        print(f"平均延迟 = {round(avg,3)} ms,标准差 = {round(std,3)} ms。")

    except KeyboardInterrupt:  
        print ("收到键盘中断信号!")
  
    except Exception as e:  
        print (f"发生其他错误或异常: {e}")  
  
    finally:  
        # 向消费者线程发送终止信号
        Global_Test_Queue.put(-1)

    return 0


if __name__ == '__main__':
    main()

树莓派示例输出

Beginning timing test for 20 seconds.
Delay times were (in mSec): [2.2789670038037, 1.2069819786120206, 1.3649799802806228, 1.4579780108761042, 1.257980999071151, 1.4789779961574823, 1.199983002152294, 1.329980994341895, 1.2979809835087508, 1.1639830190688372, 1.3999800139572471, 1.5709770086687058, 1.9739709969144315, 2.065970009425655]
Average = 1.962 ms. Standard deviaion = 0.597 ms.

优化方案与排查建议

一、性能优化方法

  1. 替换轻量通信机制
    queue.Queue依赖内部锁保证线程安全,带来额外开销。如果是单生产者单消费者场景,直接用collections.deque替代:它的append()和popleft()在GIL特性下是线程安全的,速度远快于Queue。需要自己处理阻塞逻辑时,可以配合threading.Condition手动控制等待,避免轮询浪费CPU。

  2. 降低线程切换开销

    • 绑定CPU亲和性:用os.sched_setaffinity把生产者和消费者线程绑定到同一个CPU核心,减少跨核调度的额外损耗。
    • 调整测试逻辑:原代码中time.sleep(int(TIMING_TEST_LENGTH/100))会导致sleep(0),生产者快速发送完所有消息后进入闲置,后续延迟数据可能包含线程调度休眠的额外开销。建议去掉int(),让消息均匀发送。
  3. 启用实时调度策略
    用chrt命令将测试进程设为实时调度优先级,比单纯提升进程优先级更有效:

    chrt -f 99 python your_script.py
    

二、系统配置检查

  1. 调整Python GIL切换间隔
    通过sys.setswitchinterval()减少不必要的线程切换,比如将间隔设为1ms:

    import sys
    sys.setswitchinterval(0.001)
    
  2. 安装实时内核(树莓派专属)
    Debian 11默认内核不包含PREEMPT_RT实时补丁,手动安装或编译实时内核可以大幅降低线程调度延迟,这对树莓派这类低功耗设备的实时性能提升明显。

三、性能分析工具

  1. Python内置工具
    • cProfile:分析代码各函数耗时占比,快速定位Queue相关操作的开销:
      python -m cProfile -s cumulative your_script.py
      
  2. 系统级工具
    • perf:追踪内核态和用户态的函数调用耗时,查看线程调度的具体开销:
      perf record -g python your_script.py
      perf report
      
    • pidstat:实时监控进程的上下文切换次数,判断是否因线程切换频繁导致延迟:
      pidstat -t -p <你的脚本PID> 1
      

内容的提问来源于stack exchange,提问作者jpilgrim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 20:05:31