You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django import-export导入CSV遇编码错误:charmap无法解码字节0x8d

问题:Django import-export导入CSV时编码错误

使用Django import-export导入CSV文件时,始终报错,错误信息如下:

Imported file has a wrong encoding: 'charmap' codec can't decode byte 0x8d in position 4510: character maps to

此前导入其他CSV文件一切正常,且生成该CSV时已设置编码为utf8,以下是生成CSV的Python爬取代码:

import json

import requests
from bs4 import BeautifulSoup
import pandas as pd
from csv import writer

url = "https://prowrestling.fandom.com/wiki/New_Japan_Pro_Wrestling/Roster"

page = requests.get(url)

soup = BeautifulSoup(page.content, 'html.parser')

# links = [
#     "https://prowrestling.fandom.com/" + a["href"] for a in soup.select("classname a")
# ]
links = [
    "https://prowrestling.fandom.com/" + a["href"] for a in soup.select("td a")
]
with open("real/njpw.csv", 'a', encoding="utf8", newline="") as f:
    print(f)
    wrt = writer(f)
    header = ["ring_name", "height", "weight", "born", "birth_place", "trainer", "debut", "resides"]
    wrt.writerow(header)
    for link in links:
        soup = BeautifulSoup(requests.get(link).content, "html.parser")
        ring_name = soup.h2.text.strip()
        height = soup.select_one('.pi-data-label:-soup-contains("Height") + div')
        if height is not None:
            height = height.text.strip()
        else:
            height = ""
        weight = soup.select_one('.pi-data-label:-soup-contains("Weight") + div')
        if weight is not None:
            weight = weight.text.strip()
        else:
            weight = ""
        born = soup.select_one('.pi-data-label:-soup-contains("Born") + div')
        if born is not None:
            born = born.text.strip()
        else:
            born = ""
        birth_place = soup.select_one('.pi-data-label:-soup-contains("Birth Place") + div')
        if birth_place is not None:
            birth_place = birth_place.text.strip()
        else:
            birth_place = ""
        trainer = soup.select_one('.pi-data-label:-soup-contains("Trainer") + div')
        if trainer is not None:
            trainer = trainer.text.strip()
        else:
            trainer = ""
        debut = soup.select_one('.pi-data-label:-soup-contains("Debut") + div')
        if debut is not None:
            debut = debut.text.strip()
        else:
            debut = ""
        resides = soup.select_one('.pi-data-label:-soup-contains("Resides") + div')
        if resides is not None:
            resides = resides.text.strip()
        else:
            resides = ""
        table = [ring_name, height, weight, born, birth_place, trainer, debut, resides]

        wrt.writerow(table)
        print(table)

    #gimmik = soup.select_one('.InformationBoxTitle:-soup-contains("Current gimmick") + div')
解决办法
  • 修正爬取时的编码处理:爬取单页内容时,强制指定响应编码为utf-8,避免网页返回内容编码不一致导致乱码。修改爬取循环内的代码:

    for link in links:
        response = requests.get(link)
        response.encoding = 'utf-8'  # 强制设置响应编码
        soup = BeautifulSoup(response.text, "html.parser")
        # 后续代码不变
    

    或者对提取的文本做容错处理,自动替换无法编码的字符:

    ring_name = soup.h2.text.strip().encode('utf-8', 'replace').decode('utf-8')
    # height、weight等其他字段也可按此方式处理
    
  • 增强CSV写入的容错性:打开CSV文件时添加errors='replace'参数,确保无法编码的字符被自动替换,保证文件编码严格为utf-8:

    with open("real/njpw.csv", 'a', encoding="utf8", errors='replace', newline="") as f:
    
  • 指定Django import-export的导入编码:在Resource类中显式设置编码为utf-8,确保导入时用正确编码读取文件:

    from import_export import resources
    from .models import YourModel
    
    class NJPWResource(resources.ModelResource):
        class Meta:
            model = YourModel
            encoding = 'utf-8'  # 添加编码设置
    
  • 过滤无效链接:爬取前检查链接有效性,跳过无法正常访问的页面,避免异常内容混入CSV:

    for link in links:
        response = requests.get(link)
        if response.status_code != 200:
            continue  # 跳过无效链接
        # 后续处理代码
    

内容的提问来源于stack exchange,提问作者user19353374

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 12:21:25