Python中requests获取response.text解码失败乱码问题求助
问题:Python请求网页返回乱码
我运行以下Python代码:
import requests from bs4 import BeautifulSoup url = "https://site" paramsRequest = { } headersRequest = { "Content-Type": "text/html; charset=utf-8", } response = requests.get(url, params=paramsRequest, headers=headersRequest) file_path = 'D:/tempfiles/file.txt' with open(file_path, 'w', encoding='utf-8') as file: file.write(response.text) with open(file_path, 'r', encoding='utf-8') as file: file_content = file.read() print(file_content) soup = BeautifulSoup(response.text, 'html.parser') print(soup.prettify()) response.encoding = response.apparent_encoding print(response.text)
但返回乱码结果:
�� ��9�?���/����@�[��08�/��)�#�h4�f���A��ѸQ(M{�B���������4�k(i76���r0� ;.�b>�s�$g�& 8��x�G7E��88&�s�jܦ��+�7��&)���1�������� )�~,�d���� }c j������?�$�e�T��K�K�4��c�v:�K���)�TJ �X����H5�é�ߏ�r)�:����_���
我已经尝试过两种方法:
- 设置
response.encoding = response.apparent_encoding - 使用
BeautifulSoup解析并打印soup.prettify()
补充完整请求头信息:
Accept: application/json, text/plain, */* Accept-Language:en-US;q=0.5,en;q=0.3 Accept-Encoding: gzip, deflate, br Content-Type: text/html; charset=utf-8 X-Lead-Source: website Connection: keep-alive Cookie: *ew_source*=widget Sec-Fetch-Dest: empty Sec-Fetch-Mode: cors Sec-Fetch-Site: same-origin TE: trailers answer:json
请求解决该乱码问题。
解决方案
乱码根源在于两个问题:一是请求头配置错误,二是未正确处理服务器返回的Brotli(br)压缩响应。
步骤1:修正请求头配置
GET请求无需设置Content-Type字段(该字段用于POST请求声明提交数据的类型),保留这个字段会干扰服务器返回正确的响应格式。同时,若未安装Brotli解码依赖,建议从Accept-Encoding中移除br,只保留gzip, deflate。
步骤2:调整响应处理逻辑
优先让requests自动识别编码,先设置response.encoding = response.apparent_encoding,再获取响应文本,避免编码不匹配导致的乱码。
修改后的完整代码
import requests from bs4 import BeautifulSoup url = "https://site" paramsRequest = {} headersRequest = { "Accept": "application/json, text/plain, */*", "Accept-Language": "en-US;q=0.5,en;q=0.3", "Accept-Encoding": "gzip, deflate", "X-Lead-Source": "website", "Connection": "keep-alive", "Cookie": "*ew_source*=widget", "Sec-Fetch-Dest": "empty", "Sec-Fetch-Mode": "cors", "Sec-Fetch-Site": "same-origin", "TE": "trailers", "answer": "json" } response = requests.get(url, params=paramsRequest, headers=headersRequest) # 先设置正确编码再获取文本 response.encoding = response.apparent_encoding file_content = response.text file_path = 'D:/tempfiles/file.txt' with open(file_path, 'w', encoding='utf-8') as file: file.write(file_content) print(file_content) soup = BeautifulSoup(file_content, 'html.parser') print(soup.prettify())
可选:支持Brotli压缩
如果需要保留br压缩格式,先安装Brotli依赖:
pip install brotli
安装后,requests会自动处理Brotli编码的响应,此时请求头中可以保留Accept-Encoding: gzip, deflate, br。
内容的提问来源于stack exchange,提问作者olegch
相关产品推荐
相关产品推荐

