如何用Python和Numpy复现Matlab的.dat文件读取解码逻辑?
Matlab转Python+Numpy解码.dat文件结果不一致问题
我有一段Matlab脚本可读取并解码编码后的.dat文件,再保存结果。转译为Python+Numpy实现后,同一文件的输出结果不符(Python生成数值无意义),该数据结构来自串口读取脚本。
起初怀疑是位移位问题(索引差异、Numpy区分左右移位函数),排查后排除;拆分移位与点积运算确认顺序无误,仍未解决。将Python的.read()替换为np.fromfile()并指定uint8格式后,调试发现packet_data数值与Matlab完全不同,并非reshape导致的顺序打乱,而是整个矩阵数值错误,进而packet_data_bytes也完全错误。不清楚同一文件在Python和Matlab中读取结果不同的原因,不确定是读取函数差异还是文件打开方式导致。
Matlab原代码
% Define input and output file paths inputFilePath = 'C:/Users/x/Downloads/Raw.dat' outputFilePath = 'C:/Users/x/Downloads/decoded.dat' % Open the input file fileID = fopen(inputFilePath, 'r'); % Open the output file outputFileID = fopen(outputFilePath, 'w'); % Check if files are successfully opened if fileID == -1 error('Cannot open the input file.'); end if outputFileID == -1 error('Cannot open the output file.'); end % Constants packetNumberBytesLength = 4; packetDataBytesRows = 12; packetDataBytesCols = 1024; packetDataMasks = [1,2,4,8,16,32]; numSamples = 6; numChannels = 1024; % Read and process each packet while ~feof(fileID) % Read packet number bytes packetNumberBytes = fread(fileID, packetNumberBytesLength, 'uint8'); if numel(packetNumberBytes) < packetNumberBytesLength break; end % Read packet data bytes packetDataBytes = fread(fileID, [packetDataBytesRows, packetDataBytesCols], 'uint8'); if numel(packetDataBytes) < packetDataBytesRows * packetDataBytesCols break; end % Decode packet number packetNumber = [1677216, 65536, 256, 1] * packetNumberBytes; % Decoding packet data Samples = zeros(numSamples, numChannels); for n = 1:numSamples Samples(n, :) = [2048,1024,512,256,128,64,32,16,8,4,2,1] * bitshift(bitand(packetDataBytes, packetDataMasks(n)), 1-n); Samples(n, numChannels) = 0; % Invalid sample in case of Single Cell scan mode end % Get current date and time currentDateTime = datestr(now, 'dd-mmm-yyyy HH:MM:SS.FFF'); % Write decoded data to output file writematrix(currentDateTime, outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); writematrix([packetNumber, 1], outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); writematrix(Samples(1, :), outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); writematrix([3], outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); writematrix(Samples(3, :), outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); writematrix([5], outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); writematrix(Samples(5, :), outputFilePath, 'Delimiter', 'tab', 'WriteMode', 'append'); end % Close the input and output files fclose(fileID); fclose(outputFileID);
Python转译代码(原版本)
import numpy as np import datetime # Define input and output file paths input_file_path = 'C:/Users/x/Downloads/Raw.dat' output_file_path = 'C:/Users/x/Downloads/decoded.dat' # Constants packet_number_bytes_length = 4 packet_data_bytes_rows = 12 packet_data_bytes_cols = 1024 num_samples = 6 num_channels = 1024 packet_data_masks = [1, 2, 4, 8, 16, 32] # Open the input and output files try: with open(input_file_path, 'rb') as input_file, open(output_file_path, 'w') as output_file: # Read and process each packet while True: # Read packet number bytes packet_number_bytes = np.fromfile(input_file, dtype=np.uint8, count=packet_number_bytes_length) if len(packet_number_bytes) < packet_number_bytes_length: break # Read packet data bytes packet_data_bytes = np.fromfile(input_file, dtype=np.uint8, count=packet_data_bytes_rows * packet_data_bytes_cols) if len(packet_data_bytes) < packet_data_bytes_rows * packet_data_bytes_cols: break # Reshape the packet data into a 2D array packet_data = packet_data_bytes.reshape((packet_data_bytes_rows, packet_data_bytes_cols)) # Decode packet number packet_number = np.dot([1677216, 65536, 256, 1], np.frombuffer(packet_number_bytes, dtype=np.uint8)) # Decode packet data Samples = np.zeros((num_samples, num_channels), dtype=int) for n in range(num_samples): Samples[n, :] = np.dot( [2048, 1024, 512, 256, 128, 64, 32, 16, 8, 4, 2, 1], np.right_shift(np.bitwise_and(packet_data, packet_data_masks[n]), n) ) Samples[:, num_channels - 1] = 0 # Invalid sample in case of Single Cell scan mode # Get current date and time current_date_time = datetime.datetime.now().strftime('%d-%b-%Y %H:%M:%S.%f')[:-3] # Write decoded data to output file output_file.write(current_date_time + '\n') output_file.write(f"{packet_number}\t1\n") np.savetxt(output_file, Samples[0, :].reshape(1, -1), delimiter='\t', fmt='%d') output_file.write('3\n') np.savetxt(output_file, Samples[2, :].reshape(1, -1), delimiter='\t', fmt='%d') output_file.write('5\n') np.savetxt(output_file, Samples[4, :].reshape(1, -1), delimiter='\t', fmt='%d') except FileNotFoundError as e: print(f"Error: {e}")
核心问题排查与修正方案
1. 文件打开模式差异(Matlab端)
Matlab中fopen(inputFilePath, 'r')默认是文本模式,在Windows系统下会自动转换换行符(\r\n转为\n),导致二进制数据损坏。需改为二进制模式读取:
fileID = fopen(inputFilePath, 'rb');
2. 数组维度填充顺序差异(Python端)
Matlab的fread是列优先填充数组,而Numpy默认reshape是行优先(C-style)。需指定列优先(Fortran-style)匹配Matlab行为:
packet_data = packet_data_bytes.reshape((packet_data_bytes_rows, packet_data_bytes_cols), order='F')
3. 数据包编号解码冗余(Python端)
packet_number_bytes已经是np.uint8类型,无需重复调用np.frombuffer:
packet_number = np.dot([1677216, 65536, 256, 1], packet_number_bytes)
修正后的Python代码
import numpy as np import datetime # Define input and output file paths input_file_path = 'C:/Users/x/Downloads/Raw.dat' output_file_path = 'C:/Users/x/Downloads/decoded.dat' # Constants packet_number_bytes_length = 4 packet_data_bytes_rows = 12 packet_data_bytes_cols = 1024 num_samples = 6 num_channels = 1024 packet_data_masks = [1, 2, 4, 8, 16, 32] # Open the input and output files try: with open(input_file_path, 'rb') as input_file, open(output_file_path, 'w') as output_file: # Read and process each packet while True: # Read packet number bytes packet_number_bytes = np.fromfile(input_file, dtype=np.uint8, count=packet_number_bytes_length) if len(packet_number_bytes) < packet_number_bytes_length: break # Read packet data bytes packet_data_bytes = np.fromfile(input_file, dtype=np.uint8, count=packet_data_bytes_rows * packet_data_bytes_cols) if len(packet_data_bytes) < packet_data_bytes_rows * packet_data_bytes_cols: break # Reshape using Fortran-style (column-major) to match Matlab's fread packet_data = packet_data_bytes.reshape((packet_data_bytes_rows, packet_data_bytes_cols), order='F') # Decode packet number packet_number = np.dot([1677216, 65536, 256, 1], packet_number_bytes) # Decode packet data Samples = np.zeros((num_samples, num_channels), dtype=int) for n in range(num_samples): shifted = np.right_shift(np.bitwise_and(packet_data, packet_data_masks[n]), n) Samples[n, :] = np.dot( [2048, 1024, 512, 256, 128, 64, 32, 16, 8, 4, 2, 1], shifted ) Samples[:, num_channels - 1] = 0 # Invalid sample in case of Single Cell scan mode # Get current date and time current_date_time = datetime.datetime.now().strftime('%d-%b-%Y %H:%M:%S.%f')[:-3] # Write decoded data to output file output_file.write(current_date_time + '\n') output_file.write(f"{packet_number}\t1\n") np.savetxt(output_file, Samples[0, :].reshape(1, -1), delimiter='\t', fmt='%d') output_file.write('3\n') np.savetxt(output_file, Samples[2, :].reshape(1, -1), delimiter='\t', fmt='%d') output_file.write('5\n') np.savetxt(output_file, Samples[4, :].reshape(1, -1), delimiter='\t', fmt='%d') except FileNotFoundError as e: print(f"Error: {e}")
内容的提问来源于stack exchange,提问作者gunk
相关产品推荐
相关产品推荐

