Python2.7转3.10:字节写入文件报错,如何修正?
问题修正方案
核心问题
你的代码同时混合了bytes和str两种类型:从二进制文件读取的line是bytes,处理时部分token被赋值为str(比如token = "\"\""),导致line_tokens列表中存在两种类型,delimiter.join()无法处理混合类型;同时二进制文件写入需要bytes,但"{}\n".format(...)生成的是str,双重类型冲突触发报错。
推荐采用**全程使用字符串(str)**的方式处理,Python3中字符串默认是Unicode,处理文本更直观,也能彻底避免编码混乱。
修改后的完整代码
import os def fix_csv_file(csv_file_path, delimiter="\t"): temp_file_path = "{}.temp".format(csv_file_path) # 以文本模式打开文件,指定utf-8编码 with open(csv_file_path, "r", encoding="utf-8") as source: with open(temp_file_path, "w", encoding="utf-8") as destination: for line in source: # 一键移除所有换行符,替代原有的endswith判断逻辑 line = line.rstrip("\r\n") # 直接用字符串分隔符分割列 line_tokens = line.split(delimiter) for idx, token in enumerate(line_tokens): token = token.strip() if token == "(null)" or token == "\"(null)\"": token = "\"\"" else: # 直接判断字符串的首尾,无需手动编码 if not token.startswith("\"") and not token.endswith("\""): token = "\"{}\"".format(token) line_tokens[idx] = token # 拼接后直接写入文本文件,无需转字节 destination.write("{}\n".format(delimiter.join(line_tokens))) os.remove(csv_file_path) os.rename(temp_file_path, csv_file_path)
关键修改点说明
- 文件打开模式调整:将
rb/wb改为r/w并指定encoding="utf-8",直接读写str类型,省去手动处理字节编码的麻烦。 - 换行符简化处理:用
line.rstrip("\r\n")替代原有的endswith判断+切片操作,更简洁且适配文本模式的换行符规则。 - 移除所有
.encode("utf-8"):全程使用字符串操作,无需再手动转换为字节类型。 - 类型完全统一:所有
token处理后均为str,确保delimiter.join()能正常拼接,写入文件时也无需额外类型转换。
备选方案(全程用字节类型)
如果因特殊需求必须用二进制模式处理,需确保所有操作都统一使用bytes类型:
import os def fix_csv_file(csv_file_path, delimiter="\t"): temp_file_path = "{}.temp".format(csv_file_path) delimiter_bytes = delimiter.encode("utf-8") with open(csv_file_path, "rb") as source: with open(temp_file_path, "wb") as destination: for line in source: if line.endswith(b"\r\n"): line = line[:-2] elif line.endswith(b"\n"): line = line[:-1] line_tokens = line.split(delimiter_bytes) for idx, token in enumerate(line_tokens): token = token.strip() if token == b"(null)" or token == b'"(null)"': token = b'""' else: if not token.startswith(b'"') and not token.endswith(b'"'): # 使用字节类型格式化,避免转换为字符串 token = b'"%b"' % token line_tokens[idx] = token # 拼接字节后写入,添加字节类型的换行符 destination.write(delimiter_bytes.join(line_tokens) + b"\n") os.remove(csv_file_path) os.rename(temp_file_path, csv_file_path)
此方案中所有字符串常量都用b""声明为字节类型,格式化操作也采用字节专属方式,确保全程类型无冲突。
内容的提问来源于stack exchange,提问作者Lars Skaug
相关产品推荐
相关产品推荐

