You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python向二进制文件写入数组耗时过长问题求助(Zynq RFSoC平台)

Zynq RFSoC Python写SD卡循环卡顿问题解决

我用搭载嵌入式Linux的Zynq RFSoC板卡编写Python程序控制FPGA,需要将FPGA接收的数据写入板卡SD卡的二进制文件。但发现写入数据的for循环偶尔耗时远超正常情况,程序卡顿数秒,这成为系统严重瓶颈,求解决。

问题代码片段

for i in range(int(self.pulseNumber)):

    self.dma0_recv.transfer(output_buffer0)
    self.dma1_recv.transfer(output_buffer1)
    self.dma2_recv.transfer(output_buffer2)
    self.dma3_recv.transfer(output_buffer3)
    self.dma0_recv.wait()
    self.dma1_recv.wait()
    self.dma2_recv.wait()
    self.dma3_recv.wait()

    self.data_source_mmio.write(0, 12)

    begin_writing = time.time()

    filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel0.bin"
    file_path = os.path.join(folder_name, filename)
    #np.save(file_path, output_buffer0)
    with open(file_path, 'wb') as file:
        file.write(output_buffer0.tobytes())

    filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel1.bin"
    file_path = os.path.join(folder_name, filename)
    #np.save(file_path, output_buffer1)
    with open(file_path, 'wb') as file:
        file.write(output_buffer1.tobytes())

    filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel2.bin"
    file_path = os.path.join(folder_name, filename)
    #np.save(file_path, output_buffer2)
    with open(file_path, 'wb') as file:
        file.write(output_buffer2.tobytes())

    filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel3.bin"
    file_path = os.path.join(folder_name, filename)
    #np.save(file_path, output_buffer3)
    with open(file_path, 'wb') as file:
        file.write(output_buffer3.tobytes())

    end_writing = time.time()

卡顿原因分析

  • 每次循环创建4个新文件,频繁的文件打开/关闭操作触发大量文件系统元数据写入,嵌入式SD卡IO性能有限,元数据操作极易引发卡顿。
  • 嵌入式Linux内核页缓存会批量刷写数据到SD卡,偶尔的大规模刷写会阻塞当前进程。
  • SD卡自身的垃圾回收、磨损均衡机制在后台运行时,会临时占用IO带宽导致写入延迟。

解决方案

1. 合并文件,减少IO操作次数

不要每次循环生成4个小文件,改为每个通道对应一个大文件,循环中追加写入,避免频繁创建文件、修改元数据的开销:

# 初始化阶段打开所有通道文件并预分配空间
channel_handles = []
for ch_idx in range(4):
    file_name = f"{self.experiment_name}_channel{ch_idx}.bin"
    full_path = os.path.join(folder_name, file_name)
    # 预分配总空间(假设每个脉冲数据大小为output_buffer0.nbytes)
    total_data_size = int(self.pulseNumber) * output_buffer0.nbytes
    with open(full_path, 'wb') as f:
        f.seek(total_data_size - 1)
        f.write(b'\x00')
    # 以追加模式打开文件句柄,循环中复用
    channel_handles.append(open(full_path, 'ab'))

try:
    for i in range(int(self.pulseNumber)):
        # DMA接收逻辑不变
        self.dma0_recv.transfer(output_buffer0)
        self.dma1_recv.transfer(output_buffer1)
        self.dma2_recv.transfer(output_buffer2)
        self.dma3_recv.transfer(output_buffer3)
        self.dma0_recv.wait()
        self.dma1_recv.wait()
        self.dma2_recv.wait()
        self.dma3_recv.wait()

        self.data_source_mmio.write(0, 12)

        # 直接追加写入复用的文件句柄
        channel_handles[0].write(output_buffer0.tobytes())
        channel_handles[1].write(output_buffer1.tobytes())
        channel_handles[2].write(output_buffer2.tobytes())
        channel_handles[3].write(output_buffer3.tobytes())

        # 定期刷新缓存,避免一次性刷写过多数据(比如每10次循环刷一次)
        if i % 10 == 0:
            for handle in channel_handles:
                handle.flush()
finally:
    # 最后统一关闭所有文件句柄
    for handle in channel_handles:
        handle.close()

2. 用异步线程分离IO和数据接收

把文件写入操作放到独立线程,主循环只负责接收FPGA数据并放入队列,避免IO卡顿阻塞DMA接收流程:

import threading
from queue import Queue

# 创建数据队列,限制大小防止内存溢出
write_queue = Queue(maxsize=15)

def file_write_worker(folder, exp_name, total_pulses, single_buf_size):
    # 初始化通道文件
    channel_files = []
    for ch in range(4):
        file_path = os.path.join(folder, f"{exp_name}_channel{ch}.bin")
        total_size = total_pulses * single_buf_size
        with open(file_path, 'wb') as f:
            f.seek(total_size - 1)
            f.write(b'\x00')
        channel_files.append(open(file_path, 'ab'))
    
    try:
        while True:
            data = write_queue.get()
            if data is None:  # 终止信号
                break
            buf0, buf1, buf2, buf3 = data
            channel_files[0].write(buf0.tobytes())
            channel_files[1].write(buf1.tobytes())
            channel_files[2].write(buf2.tobytes())
            channel_files[3].write(buf3.tobytes())
            write_queue.task_done()
    finally:
        for f in channel_files:
            f.close()

# 启动写入线程
buf_size = output_buffer0.nbytes
worker_thread = threading.Thread(
    target=file_write_worker,
    args=(folder_name, self.experiment_name, int(self.pulseNumber), buf_size)
)
worker_thread.start()

try:
    for i in range(int(self.pulseNumber)):
        # DMA接收逻辑
        self.dma0_recv.transfer(output_buffer0)
        self.dma1_recv.transfer(output_buffer1)
        self.dma2_recv.transfer(output_buffer2)
        self.dma3_recv.transfer(output_buffer3)
        self.dma0_recv.wait()
        self.dma1_recv.wait()
        self.dma2_recv.wait()
        self.dma3_recv.wait()

        self.data_source_mmio.write(0, 12)

        # 将数据放入队列,不阻塞主循环
        write_queue.put((output_buffer0, output_buffer1, output_buffer2, output_buffer3))
finally:
    # 发送终止信号并等待线程结束
    write_queue.put(None)
    worker_thread.join()

3. 优化SD卡挂载参数

修改/etc/fstab中SD卡的挂载选项,减少不必要的IO开销(注意:data=writeback会降低数据安全性,断电可能丢失未刷写的数据):

/dev/mmcblk0p1 /media/sd ext4 noatime,nodiratime,data=writeback 0 2

4. 更换高性能SD卡

换成UHS-I或更高等级的SD卡,提升硬件IO性能,减少底层延迟。

内容的提问来源于stack exchange,提问作者Darinoos47

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 20:40:55