You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ARM64平台共享内存IPC吞吐量低于管道与System V队列的原因及优化咨询

ARM64架构下共享内存IPC性能低于管道/System V消息队列的原因分析与优化建议

我在ARM64架构的PicoCore™MX8MP模块上开发应用时,对进程间通信(IPC)方式做了性能基准测试,对比共享内存、管道及System V消息队列的性能。令人意外的是,原本因减少数据拷贝、低开销被预期吞吐量最快的共享内存,实际性能反而低于管道和System V消息队列,特此寻求该现象的原因及优化方案。

测试细节

  • 计时方法:使用clock_gettime结合CLOCK_MONOTONIC记录性能——发送进程开始传输数据前记录起始时间戳,接收进程接收完所有数据后记录结束时间戳。
  • 共享内存测试场景:仅由发送端写入,无需同步或锁机制。
  • 瓶颈排查:已开展详细计时分析,分别捕获发送进程创建共享内存空间及数据拷贝的精确时长、接收进程挂载共享内存的耗时,各环节耗时总和与端到端测试结果一致,但整体性能未得到改善。

测试结果

自有x86-64设备

type size(MB) latency(ms) throughput(MB/sec)
pipes 1 1.700 588.2000
pipes 4 5.080 787.4000
pipes 16 20.654 774.6000
pipes 64 78.338 816.9000
pipes 256 291.878 877.0000
pipes 512 599.711 853.7000
shared 1 1.128 886.5000
shared 4 4.201 952.1000
shared 16 14.075 1136.7000
shared 64 58.787 1088.6000
shared 256 221.928 1153.5000
shared 512 395.928 1293.1000
queue 1 1.651 605.6000
queue 4 4.911 814.4000
queue 16 19.063 839.3000
queue 64 71.049 900.7000
queue 256 294.923 868.0000
queue 512 536.234 954.8000

PicoCore模块(ARM64)

type size(MB) latency(ms) throughput(MB/sec)
pipes 1 2.463 406.0000
pipes 4 6.882 581.2000
pipes 16 22.260 718.7000
pipes 64 85.234 750.8000
pipes 256 333.581 767.4000
pipes 512 665.159 769.7000
shared 1 2.795 357.7000
shared 4 9.454 423.1000
shared 16 36.306 440.6000
shared 64 141.735 451.5000
shared 256 572.561 447.1000
shared 512 1134.393 451.3000
queue 1 2.288 437.0000
queue 4 7.145 559.8000
queue 16 26.403 605.9000
queue 64 106.813 599.1000
queue 256 425.763 601.2000
queue 512 853.392 599.9000

恳请提供相关分析见解及优化建议。


内容的提问来源于stack exchange,提问作者Jeppe Lindhard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 19:47:04