Python向二进制文件写入数组耗时过长问题求助(Zynq RFSoC平台)
Zynq RFSoC Python写SD卡循环卡顿问题解决
我用搭载嵌入式Linux的Zynq RFSoC板卡编写Python程序控制FPGA,需要将FPGA接收的数据写入板卡SD卡的二进制文件。但发现写入数据的for循环偶尔耗时远超正常情况,程序卡顿数秒,这成为系统严重瓶颈,求解决。
问题代码片段
for i in range(int(self.pulseNumber)): self.dma0_recv.transfer(output_buffer0) self.dma1_recv.transfer(output_buffer1) self.dma2_recv.transfer(output_buffer2) self.dma3_recv.transfer(output_buffer3) self.dma0_recv.wait() self.dma1_recv.wait() self.dma2_recv.wait() self.dma3_recv.wait() self.data_source_mmio.write(0, 12) begin_writing = time.time() filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel0.bin" file_path = os.path.join(folder_name, filename) #np.save(file_path, output_buffer0) with open(file_path, 'wb') as file: file.write(output_buffer0.tobytes()) filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel1.bin" file_path = os.path.join(folder_name, filename) #np.save(file_path, output_buffer1) with open(file_path, 'wb') as file: file.write(output_buffer1.tobytes()) filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel2.bin" file_path = os.path.join(folder_name, filename) #np.save(file_path, output_buffer2) with open(file_path, 'wb') as file: file.write(output_buffer2.tobytes()) filename = self.experiment_name + "_waveform" + str(i).zfill(4) + "_channel3.bin" file_path = os.path.join(folder_name, filename) #np.save(file_path, output_buffer3) with open(file_path, 'wb') as file: file.write(output_buffer3.tobytes()) end_writing = time.time()
卡顿原因分析
- 每次循环创建4个新文件,频繁的文件打开/关闭操作触发大量文件系统元数据写入,嵌入式SD卡IO性能有限,元数据操作极易引发卡顿。
- 嵌入式Linux内核页缓存会批量刷写数据到SD卡,偶尔的大规模刷写会阻塞当前进程。
- SD卡自身的垃圾回收、磨损均衡机制在后台运行时,会临时占用IO带宽导致写入延迟。
解决方案
1. 合并文件,减少IO操作次数
不要每次循环生成4个小文件,改为每个通道对应一个大文件,循环中追加写入,避免频繁创建文件、修改元数据的开销:
# 初始化阶段打开所有通道文件并预分配空间 channel_handles = [] for ch_idx in range(4): file_name = f"{self.experiment_name}_channel{ch_idx}.bin" full_path = os.path.join(folder_name, file_name) # 预分配总空间(假设每个脉冲数据大小为output_buffer0.nbytes) total_data_size = int(self.pulseNumber) * output_buffer0.nbytes with open(full_path, 'wb') as f: f.seek(total_data_size - 1) f.write(b'\x00') # 以追加模式打开文件句柄,循环中复用 channel_handles.append(open(full_path, 'ab')) try: for i in range(int(self.pulseNumber)): # DMA接收逻辑不变 self.dma0_recv.transfer(output_buffer0) self.dma1_recv.transfer(output_buffer1) self.dma2_recv.transfer(output_buffer2) self.dma3_recv.transfer(output_buffer3) self.dma0_recv.wait() self.dma1_recv.wait() self.dma2_recv.wait() self.dma3_recv.wait() self.data_source_mmio.write(0, 12) # 直接追加写入复用的文件句柄 channel_handles[0].write(output_buffer0.tobytes()) channel_handles[1].write(output_buffer1.tobytes()) channel_handles[2].write(output_buffer2.tobytes()) channel_handles[3].write(output_buffer3.tobytes()) # 定期刷新缓存,避免一次性刷写过多数据(比如每10次循环刷一次) if i % 10 == 0: for handle in channel_handles: handle.flush() finally: # 最后统一关闭所有文件句柄 for handle in channel_handles: handle.close()
2. 用异步线程分离IO和数据接收
把文件写入操作放到独立线程,主循环只负责接收FPGA数据并放入队列,避免IO卡顿阻塞DMA接收流程:
import threading from queue import Queue # 创建数据队列,限制大小防止内存溢出 write_queue = Queue(maxsize=15) def file_write_worker(folder, exp_name, total_pulses, single_buf_size): # 初始化通道文件 channel_files = [] for ch in range(4): file_path = os.path.join(folder, f"{exp_name}_channel{ch}.bin") total_size = total_pulses * single_buf_size with open(file_path, 'wb') as f: f.seek(total_size - 1) f.write(b'\x00') channel_files.append(open(file_path, 'ab')) try: while True: data = write_queue.get() if data is None: # 终止信号 break buf0, buf1, buf2, buf3 = data channel_files[0].write(buf0.tobytes()) channel_files[1].write(buf1.tobytes()) channel_files[2].write(buf2.tobytes()) channel_files[3].write(buf3.tobytes()) write_queue.task_done() finally: for f in channel_files: f.close() # 启动写入线程 buf_size = output_buffer0.nbytes worker_thread = threading.Thread( target=file_write_worker, args=(folder_name, self.experiment_name, int(self.pulseNumber), buf_size) ) worker_thread.start() try: for i in range(int(self.pulseNumber)): # DMA接收逻辑 self.dma0_recv.transfer(output_buffer0) self.dma1_recv.transfer(output_buffer1) self.dma2_recv.transfer(output_buffer2) self.dma3_recv.transfer(output_buffer3) self.dma0_recv.wait() self.dma1_recv.wait() self.dma2_recv.wait() self.dma3_recv.wait() self.data_source_mmio.write(0, 12) # 将数据放入队列,不阻塞主循环 write_queue.put((output_buffer0, output_buffer1, output_buffer2, output_buffer3)) finally: # 发送终止信号并等待线程结束 write_queue.put(None) worker_thread.join()
3. 优化SD卡挂载参数
修改/etc/fstab中SD卡的挂载选项,减少不必要的IO开销(注意:data=writeback会降低数据安全性,断电可能丢失未刷写的数据):
/dev/mmcblk0p1 /media/sd ext4 noatime,nodiratime,data=writeback 0 2
4. 更换高性能SD卡
换成UHS-I或更高等级的SD卡,提升硬件IO性能,减少底层延迟。
内容的提问来源于stack exchange,提问作者Darinoos47
相关产品推荐
相关产品推荐

