You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取含德语变音符号CSV报错_csv.Error: line contains NUL问题咨询

报错原因说明

你遇到的_csv.Error: line contains NUL报错核心原因是读取文件时未指定匹配的编码格式:你的CSV为UTF-16 LE编码,Windows环境下Python默认用GBK编码读取多字节编码的文件,会错误解析出大量零值字节(NUL),导致csv模块无法正常处理。你之前尝试了各类编码但未在open函数中显式指定,所以不生效。


方案1:修改现有CSV处理脚本

直接适配你现有的UTF-16 LE编码CSV文件,修改后的代码如下:

import csv

# 读取时指定编码匹配你的文件,同时提前清理NUL字符
with open("products3.csv", newline='', encoding='utf-16-le') as r_file:
    # 替换所有NUL字符后再按行拆分
    content = r_file.read().replace('\0', '')
    file_reader = csv.reader(content.splitlines(), delimiter = ";")
    # 写入指定utf-8-sig编码,Excel打开不会出现德语变音符号乱码
    with open("products_new.csv", mode="w", newline='', encoding='utf-8-sig') as w_file:
        file_writer = csv.writer(w_file, delimiter = ";")
        count = 0
        for row in file_reader:
            # 跳过空行避免索引报错
            if not row:
                continue
            if row[24].find(';') != -1:
                img_srcs =  row[24].split(';')
                handle = row[0]
                # 清理图片链接前后多余空格
                row[24] = img_srcs[0].strip()
                file_writer.writerow(row)
                index = 0
                for img in img_srcs:
                    img = img.strip()
                    if img:
                        if index != 0:
                            # 动态生成空列,避免手动拼写错误
                            new_row = [handle] + ['']*24 + [img] + ['']*(len(row)-25)
                            file_writer.writerow(new_row)
                    index += 1
            else:
                file_writer.writerow(row)
            count += 1
            if count == 3158:
                break
        print(f'Total products in file: {count}')

关键改动说明:

  • 读取时显式指定encoding='utf-16-le'匹配你的文件编码
  • 提前替换内容中的NUL字符,避免csv模块解析报错
  • 写入时使用utf-8-sig编码,输出的CSV可直接用Excel打开,不会出现德语变音符号显示为问号的问题
  • 删除多余的手动关闭文件代码,with语法会自动管理文件句柄
  • 增加空行判断和空格清理逻辑,避免异常空值报错

方案2:直接读取原始XLSX文件(更推荐)

完全避开文件格式转换的编码问题,直接处理原始.xlsx文件,不需要手动转CSV:

  1. 先安装依赖库:
    pip install pandas openpyxl
  2. 处理代码如下:
import pandas as pd

# 直接读取原始xlsx文件,自动处理所有特殊字符
df = pd.read_excel("products3.xlsx", dtype=str)
result_rows = []

for _, row in df.iterrows():
    img_src = str(row.iloc[24]).strip()
    if ';' in img_src:
        # 拆分并清理多余的空图片链接
        img_list = [i.strip() for i in img_src.split(';') if i.strip()]
        # 第一行保留原数据,替换为第一张图片
        row.iloc[24] = img_list[0]
        result_rows.append(row.copy())
        # 剩余图片单独生成新行
        for img in img_list[1:]:
            new_row = pd.Series(['']*len(df.columns), index=df.columns)
            new_row.iloc[0] = row.iloc[0]
            new_row.iloc[24] = img
            result_rows.append(new_row)
    else:
        result_rows.append(row)

# 输出为Excel可正常打开的UTF-8编码CSV
result_df = pd.DataFrame(result_rows)
result_df.to_csv("products_new.csv", sep=";", index=False, encoding='utf-8-sig')
print(f'Total products in file: {len(result_df)}')

内容的提问来源于stack exchange,提问作者Max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 11:06:05