如何将CSV文件中的NULL字节替换为'0'?Python代码调试求助
如何将CSV文件中的NULL字节替换为'0'
我有一个含NULL字节的CSV文件,想打开后把其中的NULL字节替换成'0',但写的代码都跑不起来:
原代码:
try: # open source file with open (dataFile,'r')as csvfile: sourceDF = csv.reader(csvfile) replaced = [sourceDF.replace(b'\0',b'0') for sourceDF in replaced] print(replaced) first_line = True selHeaders = [] # read each row in source file for dataRow in sourceDF: # check if first line of file if first_line == True: first_line = False first_row = dataRow # check if first file in compile list if first_run == 0: result.append(list())
还试过这段代码,同样没用:
replaced = [sourceDF.replace(b'\0',b'0') for sourceDF in replaced] print(replaced)
问题出在哪
- 文本/字节模式不匹配:用
'r'文本模式打开文件,操作的是字符串,但你用了字节类型的b'\0',两者没法直接替换。 - 逻辑顺序错误:
csv.reader是迭代器,还没读取内容就想调用replace,而且列表推导里的replaced变量未定义就使用,直接会触发报错。 - 迭代器无法直接调用replace:
csv.reader返回的是逐行迭代的对象,不是字符串类型,不能直接用replace方法。
正确的处理方法
方案1:一次性处理(适合小文件)
先把整个文件读入内存,替换所有NULL字节后再解析CSV:
import csv from io import StringIO dataFile = "你的CSV文件路径" result = [] first_run = 0 try: # 二进制模式读文件,替换NULL字节后转成字符串 with open(dataFile, 'rb') as f: content = f.read().replace(b'\0', b'0').decode('utf-8') # 用StringIO模拟文件对象,交给csv.reader解析 with StringIO(content) as csvfile: reader = csv.reader(csvfile) first_line = True selHeaders = [] for data_row in reader: if first_line: first_line = False first_row = data_row if first_run == 0: result.append([]) # 此处添加你的后续处理逻辑 print(data_row) except Exception as e: print(f"出错了:{e}")
方案2:逐行处理(适合大文件)
如果文件体积过大,不想一次性占用过多内存,可逐行读取替换:
import csv from io import StringIO dataFile = "你的CSV文件路径" result = [] first_run = 0 try: with open(dataFile, 'rb') as f: line_num = 0 for line in f: line_num += 1 # 替换当前行的NULL字节,转成字符串 processed_line = line.replace(b'\0', b'0').decode('utf-8') # 解析当前行数据 row = next(csv.reader(StringIO(processed_line))) if line_num == 1: first_row = row if first_run == 0: result.append([]) # 此处添加你的后续处理逻辑 print(row) except Exception as e: print(f"出错了:{e}")
内容的提问来源于stack exchange,提问作者Zhang
相关产品推荐
相关产品推荐

