You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中requests获取response.text解码失败乱码问题求助

问题:Python请求网页返回乱码

我运行以下Python代码:

import requests
from bs4 import BeautifulSoup

url = "https://site"

paramsRequest = {

}

headersRequest = {
"Content-Type": "text/html; charset=utf-8",
}

response = requests.get(url, params=paramsRequest, headers=headersRequest)

file_path = 'D:/tempfiles/file.txt'

with open(file_path, 'w', encoding='utf-8') as file:
    file.write(response.text)

with open(file_path, 'r', encoding='utf-8') as file:
    file_content = file.read()

print(file_content)
soup = BeautifulSoup(response.text, 'html.parser') 
print(soup.prettify())

response.encoding = response.apparent_encoding
print(response.text)

但返回乱码结果:

�� ��9�?���/����@�[��08�/��)�#�h4�f���A��ѸQ(M{�B���������4�k(i76���r0�
;.�b>�s�$g�& 8��x�G7E��88&�s�jܦ��+�7��&)���1�������� )�~,�d���� }c
j������?�$�e�T��K�K�4��c�v:�K���)�TJ
�X����H5�é�ߏ�r)�:����_���

我已经尝试过两种方法:

  • 设置response.encoding = response.apparent_encoding
  • 使用BeautifulSoup解析并打印soup.prettify()

补充完整请求头信息:

Accept: application/json, text/plain, */*
Accept-Language:en-US;q=0.5,en;q=0.3
Accept-Encoding: gzip, deflate, br
Content-Type: text/html; charset=utf-8
X-Lead-Source: website
Connection: keep-alive
Cookie: *ew_source*=widget
Sec-Fetch-Dest: empty
Sec-Fetch-Mode: cors
Sec-Fetch-Site: same-origin
TE: trailers
answer:json

请求解决该乱码问题。


解决方案

乱码根源在于两个问题:一是请求头配置错误,二是未正确处理服务器返回的Brotli(br)压缩响应。

步骤1:修正请求头配置

GET请求无需设置Content-Type字段(该字段用于POST请求声明提交数据的类型),保留这个字段会干扰服务器返回正确的响应格式。同时,若未安装Brotli解码依赖,建议从Accept-Encoding中移除br,只保留gzip, deflate。

步骤2:调整响应处理逻辑

优先让requests自动识别编码,先设置response.encoding = response.apparent_encoding,再获取响应文本,避免编码不匹配导致的乱码。

修改后的完整代码

import requests
from bs4 import BeautifulSoup

url = "https://site"

paramsRequest = {}

headersRequest = {
    "Accept": "application/json, text/plain, */*",
    "Accept-Language": "en-US;q=0.5,en;q=0.3",
    "Accept-Encoding": "gzip, deflate",
    "X-Lead-Source": "website",
    "Connection": "keep-alive",
    "Cookie": "*ew_source*=widget",
    "Sec-Fetch-Dest": "empty",
    "Sec-Fetch-Mode": "cors",
    "Sec-Fetch-Site": "same-origin",
    "TE": "trailers",
    "answer": "json"
}

response = requests.get(url, params=paramsRequest, headers=headersRequest)
# 先设置正确编码再获取文本
response.encoding = response.apparent_encoding
file_content = response.text

file_path = 'D:/tempfiles/file.txt'
with open(file_path, 'w', encoding='utf-8') as file:
    file.write(file_content)

print(file_content)
soup = BeautifulSoup(file_content, 'html.parser') 
print(soup.prettify())

可选:支持Brotli压缩

如果需要保留br压缩格式,先安装Brotli依赖:

pip install brotli

安装后,requests会自动处理Brotli编码的响应,此时请求头中可以保留Accept-Encoding: gzip, deflate, br。

内容的提问来源于stack exchange,提问作者olegch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 14:35:58