ARM64平台共享内存IPC吞吐量低于管道与System V队列的原因及优化咨询
ARM64架构下共享内存IPC性能低于管道/System V消息队列的原因分析与优化建议
我在ARM64架构的PicoCore™MX8MP模块上开发应用时,对进程间通信(IPC)方式做了性能基准测试,对比共享内存、管道及System V消息队列的性能。令人意外的是,原本因减少数据拷贝、低开销被预期吞吐量最快的共享内存,实际性能反而低于管道和System V消息队列,特此寻求该现象的原因及优化方案。
测试细节
- 计时方法:使用
clock_gettime结合CLOCK_MONOTONIC记录性能——发送进程开始传输数据前记录起始时间戳,接收进程接收完所有数据后记录结束时间戳。 - 共享内存测试场景:仅由发送端写入,无需同步或锁机制。
- 瓶颈排查:已开展详细计时分析,分别捕获发送进程创建共享内存空间及数据拷贝的精确时长、接收进程挂载共享内存的耗时,各环节耗时总和与端到端测试结果一致,但整体性能未得到改善。
测试结果
自有x86-64设备
type size(MB) latency(ms) throughput(MB/sec) pipes 1 1.700 588.2000 pipes 4 5.080 787.4000 pipes 16 20.654 774.6000 pipes 64 78.338 816.9000 pipes 256 291.878 877.0000 pipes 512 599.711 853.7000 shared 1 1.128 886.5000 shared 4 4.201 952.1000 shared 16 14.075 1136.7000 shared 64 58.787 1088.6000 shared 256 221.928 1153.5000 shared 512 395.928 1293.1000 queue 1 1.651 605.6000 queue 4 4.911 814.4000 queue 16 19.063 839.3000 queue 64 71.049 900.7000 queue 256 294.923 868.0000 queue 512 536.234 954.8000
PicoCore模块(ARM64)
type size(MB) latency(ms) throughput(MB/sec) pipes 1 2.463 406.0000 pipes 4 6.882 581.2000 pipes 16 22.260 718.7000 pipes 64 85.234 750.8000 pipes 256 333.581 767.4000 pipes 512 665.159 769.7000 shared 1 2.795 357.7000 shared 4 9.454 423.1000 shared 16 36.306 440.6000 shared 64 141.735 451.5000 shared 256 572.561 447.1000 shared 512 1134.393 451.3000 queue 1 2.288 437.0000 queue 4 7.145 559.8000 queue 16 26.403 605.9000 queue 64 106.813 599.1000 queue 256 425.763 601.2000 queue 512 853.392 599.9000
恳请提供相关分析见解及优化建议。
内容的提问来源于stack exchange,提问作者Jeppe Lindhard
相关产品推荐
相关产品推荐

