Django import-export导入CSV遇编码错误:charmap无法解码字节0x8d
问题:Django import-export导入CSV时编码错误
使用Django import-export导入CSV文件时,始终报错,错误信息如下:
Imported file has a wrong encoding: 'charmap' codec can't decode byte 0x8d in position 4510: character maps to
此前导入其他CSV文件一切正常,且生成该CSV时已设置编码为utf8,以下是生成CSV的Python爬取代码:
import json import requests from bs4 import BeautifulSoup import pandas as pd from csv import writer url = "https://prowrestling.fandom.com/wiki/New_Japan_Pro_Wrestling/Roster" page = requests.get(url) soup = BeautifulSoup(page.content, 'html.parser') # links = [ # "https://prowrestling.fandom.com/" + a["href"] for a in soup.select("classname a") # ] links = [ "https://prowrestling.fandom.com/" + a["href"] for a in soup.select("td a") ] with open("real/njpw.csv", 'a', encoding="utf8", newline="") as f: print(f) wrt = writer(f) header = ["ring_name", "height", "weight", "born", "birth_place", "trainer", "debut", "resides"] wrt.writerow(header) for link in links: soup = BeautifulSoup(requests.get(link).content, "html.parser") ring_name = soup.h2.text.strip() height = soup.select_one('.pi-data-label:-soup-contains("Height") + div') if height is not None: height = height.text.strip() else: height = "" weight = soup.select_one('.pi-data-label:-soup-contains("Weight") + div') if weight is not None: weight = weight.text.strip() else: weight = "" born = soup.select_one('.pi-data-label:-soup-contains("Born") + div') if born is not None: born = born.text.strip() else: born = "" birth_place = soup.select_one('.pi-data-label:-soup-contains("Birth Place") + div') if birth_place is not None: birth_place = birth_place.text.strip() else: birth_place = "" trainer = soup.select_one('.pi-data-label:-soup-contains("Trainer") + div') if trainer is not None: trainer = trainer.text.strip() else: trainer = "" debut = soup.select_one('.pi-data-label:-soup-contains("Debut") + div') if debut is not None: debut = debut.text.strip() else: debut = "" resides = soup.select_one('.pi-data-label:-soup-contains("Resides") + div') if resides is not None: resides = resides.text.strip() else: resides = "" table = [ring_name, height, weight, born, birth_place, trainer, debut, resides] wrt.writerow(table) print(table) #gimmik = soup.select_one('.InformationBoxTitle:-soup-contains("Current gimmick") + div')
解决办法
修正爬取时的编码处理:爬取单页内容时,强制指定响应编码为utf-8,避免网页返回内容编码不一致导致乱码。修改爬取循环内的代码:
for link in links: response = requests.get(link) response.encoding = 'utf-8' # 强制设置响应编码 soup = BeautifulSoup(response.text, "html.parser") # 后续代码不变或者对提取的文本做容错处理,自动替换无法编码的字符:
ring_name = soup.h2.text.strip().encode('utf-8', 'replace').decode('utf-8') # height、weight等其他字段也可按此方式处理增强CSV写入的容错性:打开CSV文件时添加
errors='replace'参数,确保无法编码的字符被自动替换,保证文件编码严格为utf-8:with open("real/njpw.csv", 'a', encoding="utf8", errors='replace', newline="") as f:指定Django import-export的导入编码:在Resource类中显式设置编码为utf-8,确保导入时用正确编码读取文件:
from import_export import resources from .models import YourModel class NJPWResource(resources.ModelResource): class Meta: model = YourModel encoding = 'utf-8' # 添加编码设置过滤无效链接:爬取前检查链接有效性,跳过无法正常访问的页面,避免异常内容混入CSV:
for link in links: response = requests.get(link) if response.status_code != 200: continue # 跳过无效链接 # 后续处理代码
内容的提问来源于stack exchange,提问作者user19353374
相关产品推荐
相关产品推荐

