批量转置.lvm文件列转行时遇UnicodeDecodeError问题求助
解决LVM文件转置时的Unicode编码错误及代码完善
一、编码错误的解决方法
你遇到的UnicodeDecodeError是因为Python默认用UTF-8编码读取文件,但你的LVM文件采用了其他编码(比如Latin-1,报错中的0xd6是该编码下的字符Ö)。解决方法有两种:
指定正确编码打开文件:如果确定文件是Latin-1编码,修改文件打开代码:
f = open(os.path.join(lvm_directory, filename), encoding='latin-1')忽略编码错误(应急方案):不确定编码时,可让Python忽略无法解码的字符:
f = open(os.path.join(lvm_directory, filename), errors='ignore')
二、完善转置逻辑的完整代码
原代码存在逻辑漏洞(仅收集表头未处理数据、路径拼接不安全),以下是修复后可直接运行的完整代码,实现将每个LVM文件的列数据转置为行并合并到输出文件:
import os # 替换为你的实际路径 lvm_directory = "/Users/**path**" output_file = "/Users/**path**/Transposed.lvm" output_delim = "\t" # 存储所有文件的转置数据 all_transposed_data = [] # 存储表头(如果所有文件表头一致,可只取第一个文件的) header = [] for filename in os.listdir(lvm_directory): # 仅处理.lvm文件 if not filename.endswith('.lvm'): continue file_path = os.path.join(lvm_directory, filename) file_data = [] # 用Latin-1编码打开文件,避免编码错误 with open(file_path, encoding='latin-1') as f: for line_num, line in enumerate(f): # 跳过空行 line = line.strip() if not line: continue sl = line.split() # 第一行作为表头(如果所有文件表头相同) if line_num == 0: if not header: header = sl continue # 收集每行的列数据(这里假设你要转置所有列,可根据需求调整) file_data.append(sl) # 将当前文件的列转置为行 if file_data: # zip(*file_data)实现转置,再转换为字符串 transposed_rows = [output_delim.join(row) for row in zip(*file_data)] # 给转置后的每行加上原文件名标识(可选) transposed_with_name = [f"{filename}\t{row}" for row in transposed_rows] all_transposed_data.extend(transposed_with_name) # 将结果写入输出文件 with open(output_file, 'w', encoding='utf-8') as f: # 写入表头(可选,若不需要可注释) f.write(output_delim.join(header) + "\n") # 写入所有转置后的数据 f.write("\n".join(all_transposed_data))
三、代码说明
os.path.join:避免手动拼接路径时出现的斜杠问题,适配Mac系统路径格式。with open(...):自动关闭文件,比手动f.close()更安全。zip(*file_data):核心转置操作,将二维列表的行和列互换。- 增加了
.lvm文件过滤:避免处理目录中的其他无关文件。 - 可按需调整:如果只需要转置某一列(比如第二列),可修改
file_data.append(sl[1]),然后转置逻辑简化为直接将该列数据拼成一行。
四、适配原始代码框架的版本(沿用朋友思路)
如果你想保留朋友的代码逻辑,修复编码错误并补全后的代码如下:
import os lvm_directory = "/Users/**path**" output_file = "/Users/**path**/Transposed.lvm" newFile = True output_delim = "\t" for filename in os.listdir(lvm_directory): if not filename.endswith('.lvm'): continue header = [] data = [] # 修复编码问题 with open(os.path.join(lvm_directory, filename), encoding='latin-1') as f: for l in f: sl = l.split() if not sl: continue # 收集表头(仅第一个文件的每行第二列作为表头) if newFile: header.append(sl[1]) # 收集每行的第二列数据 data.append(sl[1]) # 写入输出文件 with open(output_file, 'w' if newFile else 'a', encoding='utf-8') as f: if newFile: f.write(output_delim.join(header) + "\n") newFile = False f.write(output_delim.join(data) + "\n")
这个版本会把第一个文件的每行第二列作为输出表头,后续每个文件的每行第二列作为输出的一行。
内容的提问来源于stack exchange,提问作者2Black_Cats
相关产品推荐
相关产品推荐

