You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法同时获取CSV文件列表及表头,遇UnicodeDecodeError编码错误

解决Django REST Framework CSV接口的UnicodeDecodeError问题

问题核心

  • 触发UnicodeDecodeError是因为强制用utf-8读取非UTF-8编码的CSV文件(比如GBK、CP1252这类格式),字节0xa4不属于UTF-8的有效起始字节。
  • rb二进制模式不能指定encoding参数,因为二进制模式直接操作字节流,不需要编码转换。
  • chardet检测失败时,可以尝试覆盖常见的非UTF-8编码格式做兼容。

解决方案

1. 安装依赖(若未安装)

pip install chardet

2. 修改视图代码,加入编码检测与容错逻辑

替换原有的文件读取部分,新增编码自动适配逻辑:

from rest_framework import status
from rest_framework.response import Response
from rest_framework.views import APIView
from .models import CSVFile
import csv
import chardet

class CSVFileListView(APIView):
    def get(self, request, format=None):
        files = CSVFile.objects.all()
        file_data = []
        for file in files:
            columns = self.extract_columns(file)
            file_data.append({
                'id': file.id,
                'filename': file.file.name,
                'uploaded_at': file.uploaded_at,
                'columns': columns
            })
        return Response(file_data, status=status.HTTP_200_OK)

    def extract_columns(self, file):
        # 先读二进制内容做编码检测
        with open(file.file.path, 'rb') as f:
            raw_data = f.read(1024)
            detect_result = chardet.detect(raw_data)
            encoding = detect_result['encoding'] or 'utf-8'
            
            # 检测置信度低或无结果时,尝试常见编码
            if not encoding or detect_result['confidence'] < 0.7:
                common_encodings = ['gbk', 'cp1252', 'iso-8859-1', 'gb2312']
                for enc in common_encodings:
                    try:
                        with open(file.file.path, 'r', encoding=enc) as csv_file:
                            return next(csv.reader(csv_file), [])
                    except UnicodeDecodeError:
                        continue
                # 所有编码尝试失败,用容错模式读取
                with open(file.file.path, 'r', encoding='utf-8', errors='replace') as csv_file:
                    return next(csv.reader(csv_file), [])
            
            # 编码检测可信时直接读取
            with open(file.file.path, 'r', encoding=encoding) as csv_file:
                return next(csv.reader(csv_file), [])

class CSVFileDetailView(APIView):
    def get(self, request, file_id, format=None):
        try:
            file = CSVFile.objects.get(id=file_id)
            csv_data = []
            
            # 编码检测逻辑
            with open(file.file.path, 'rb') as f:
                raw_data = f.read(1024)
                detect_result = chardet.detect(raw_data)
                encoding = detect_result['encoding'] or 'utf-8'
                
                if not encoding or detect_result['confidence'] < 0.7:
                    common_encodings = ['gbk', 'cp1252', 'iso-8859-1', 'gb2312']
                    for enc in common_encodings:
                        try:
                            with open(file.file.path, 'r', encoding=enc) as csv_file:
                                dialect = csv.Sniffer().sniff(csv_file.read(1024))
                                csv_file.seek(0)
                                csv_data = list(csv.reader(csv_file, dialect=dialect))
                            break
                        except (UnicodeDecodeError, csv.Error):
                            continue
                    else:
                        # 容错读取
                        with open(file.file.path, 'r', encoding='utf-8', errors='replace') as csv_file:
                            dialect = csv.Sniffer().sniff(csv_file.read(1024))
                            csv_file.seek(0)
                            csv_data = list(csv.reader(csv_file, dialect=dialect))
                else:
                    with open(file.file.path, 'r', encoding=encoding) as csv_file:
                        dialect = csv.Sniffer().sniff(csv_file.read(1024))
                        csv_file.seek(0)
                        csv_data = list(csv.reader(csv_file, dialect=dialect))

            columns = csv_data[0] if csv_data else []
            data = csv_data[1:] if len(csv_data) > 1 else []

            return Response({
                'columns': columns,
                'data': data
            }, status=status.HTTP_200_OK)
        except CSVFile.DoesNotExist:
            return Response({'error': 'File not found'}, status=status.HTTP_404_NOT_FOUND)

3. 关键优化说明

  • 用rb模式读取二进制内容做编码检测,避免直接指定UTF-8导致解码失败。
  • 编码检测置信度不足时,遍历常见编码尝试读取,覆盖绝大多数非UTF-8场景。
  • 最后加入errors='replace'容错机制,确保接口不会因编码完全不匹配而崩溃,仅替换无法解码的字符。
  • 修复了rb模式不能加encoding参数的错误,二进制模式仅用于编码检测,实际读取用文本模式。

内容的提问来源于stack exchange,提问作者cClonner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 00:11:01