You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python2.7转3.10:字节写入文件报错,如何修正?

问题修正方案

核心问题

你的代码同时混合了bytes和str两种类型:从二进制文件读取的line是bytes,处理时部分token被赋值为str(比如token = "\"\""),导致line_tokens列表中存在两种类型,delimiter.join()无法处理混合类型;同时二进制文件写入需要bytes,但"{}\n".format(...)生成的是str,双重类型冲突触发报错。

推荐采用**全程使用字符串(str)**的方式处理,Python3中字符串默认是Unicode,处理文本更直观,也能彻底避免编码混乱。

修改后的完整代码

import os

def fix_csv_file(csv_file_path, delimiter="\t"):
    temp_file_path = "{}.temp".format(csv_file_path)
    # 以文本模式打开文件,指定utf-8编码
    with open(csv_file_path, "r", encoding="utf-8") as source:
        with open(temp_file_path, "w", encoding="utf-8") as destination:
            for line in source:
                # 一键移除所有换行符,替代原有的endswith判断逻辑
                line = line.rstrip("\r\n")
                # 直接用字符串分隔符分割列
                line_tokens = line.split(delimiter)
                for idx, token in enumerate(line_tokens):
                    token = token.strip()
                    if token == "(null)" or token == "\"(null)\"":
                        token = "\"\""
                    else:
                        # 直接判断字符串的首尾,无需手动编码
                        if not token.startswith("\"") and not token.endswith("\""):
                            token = "\"{}\"".format(token)
                    line_tokens[idx] = token
                # 拼接后直接写入文本文件,无需转字节
                destination.write("{}\n".format(delimiter.join(line_tokens)))
    os.remove(csv_file_path)
    os.rename(temp_file_path, csv_file_path)

关键修改点说明

  1. 文件打开模式调整:将rb/wb改为r/w并指定encoding="utf-8",直接读写str类型,省去手动处理字节编码的麻烦。
  2. 换行符简化处理:用line.rstrip("\r\n")替代原有的endswith判断+切片操作,更简洁且适配文本模式的换行符规则。
  3. 移除所有.encode("utf-8"):全程使用字符串操作,无需再手动转换为字节类型。
  4. 类型完全统一:所有token处理后均为str,确保delimiter.join()能正常拼接,写入文件时也无需额外类型转换。

备选方案(全程用字节类型)

如果因特殊需求必须用二进制模式处理,需确保所有操作都统一使用bytes类型:

import os

def fix_csv_file(csv_file_path, delimiter="\t"):
    temp_file_path = "{}.temp".format(csv_file_path)
    delimiter_bytes = delimiter.encode("utf-8")
    with open(csv_file_path, "rb") as source:
        with open(temp_file_path, "wb") as destination:
            for line in source:
                if line.endswith(b"\r\n"):
                    line = line[:-2]
                elif line.endswith(b"\n"):
                    line = line[:-1]
                line_tokens = line.split(delimiter_bytes)
                for idx, token in enumerate(line_tokens):
                    token = token.strip()
                    if token == b"(null)" or token == b'"(null)"':
                        token = b'""'
                    else:
                        if not token.startswith(b'"') and not token.endswith(b'"'):
                            # 使用字节类型格式化,避免转换为字符串
                            token = b'"%b"' % token
                    line_tokens[idx] = token
                # 拼接字节后写入,添加字节类型的换行符
                destination.write(delimiter_bytes.join(line_tokens) + b"\n")
    os.remove(csv_file_path)
    os.rename(temp_file_path, csv_file_path)

此方案中所有字符串常量都用b""声明为字节类型,格式化操作也采用字节专属方式,确保全程类型无冲突。

内容的提问来源于stack exchange,提问作者Lars Skaug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:18:19