You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取数据遇乱码求助:chardet库无法解决F12 Network响应乱码问题

解决爬取数据时Network响应乱码的方案
  • 处理响应压缩编码:多数网站会用gzip/deflate压缩响应内容,直接读取会出现乱码。先通过响应头的Content-Encoding判断压缩类型,再解压解码:
import gzip
from io import BytesIO
import requests

target_url = "你的目标网页URL"
resp = requests.get(target_url, headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"})

if resp.headers.get("Content-Encoding") == "gzip":
    with gzip.GzipFile(fileobj=BytesIO(resp.content)) as f:
        content = f.read().decode("utf-8")  # 替换为网页实际编码
elif resp.headers.get("Content-Encoding") == "deflate":
    import zlib
    content = zlib.decompress(resp.content).decode("utf-8")
else:
    content = resp.content.decode("utf-8")
  • 手动指定网页编码:不要依赖chardet,直接查看网页源码里的<meta charset="xxx">标签,或Network面板中Response Headers的Content-Type字段(比如text/html; charset=gbk),用该编码解码:
# 示例:网页编码为gbk时
content = resp.content.decode("gbk")
  • 模拟浏览器请求头:部分网站会根据请求头返回不同编码内容,添加完整浏览器请求头后再请求:
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3"
}
resp = requests.get(target_url, headers=headers)
content = resp.content.decode("utf-8")  # 替换为网页实际编码
  • 尝试多种常见编码:不确定编码时,依次测试utf-8、gbk、gb2312、big5、utf-16等常见编码,找到能正常显示的:
encodings = ["utf-8", "gbk", "gb2312", "big5", "utf-16"]
for enc in encodings:
    try:
        content = resp.content.decode(enc)
        print(f"成功用{enc}解码:\n{content[:500]}")
        break
    except UnicodeDecodeError:
        continue

内容的提问来源于stack exchange,提问作者volcano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 04:45:43