使用Python wget下载XML正常,但TXT/HTML文件编码异常无法读取
问题分析与解决方案
核心原因推测
你遇到的问题大概率是以下两种情况之一:
- TXT/HTML文件被服务器压缩传输:服务器对TXT、HTML返回了gzip/deflate压缩后的二进制内容,但
wget或urllib未自动解压,直接保存的压缩文件自然会显示乱码;而XML文件可能未被服务器压缩,因此正常。 - 编码不匹配:服务器返回的TXT/HTML使用了非UTF-8编码(如GBK、GB2312、UTF-16等),但你强制用UTF-8读取,导致编码解析错误。
验证方法
先通过简单代码排查问题根源:
import requests # 替换成你的TXT/HTML链接 url = "https://.../urllist.txt" resp = requests.get(url) # 打印响应头,重点看Content-Encoding和Content-Type print("响应头:", resp.headers) # 打印自动推断的编码 print("推断编码:", resp.encoding)
- 如果
Content-Encoding显示gzip或deflate,说明是压缩问题; - 如果
Content-Type里的charset不是utf-8,说明是编码不匹配问题。
解决方案:使用Requests库处理(推荐)
requests库会自动处理gzip/deflate压缩,并能从响应头或内容中自动推断正确编码,完美解决这两类问题。修改后的代码如下:
import os import requests from django.conf import settings sitemaps = [ "https://.../sitemap.xml", "https://.../sitemap_images.xml", "https://.../sitemap_video.xml", "https://.../sitemap_mobile.xml", "https://.../sitemap.html", "https://.../urllist.txt", "https://.../ror.xml" ] def download_and_save(url): save_dir = settings.STATICFILES_DIRS[0] filename = url.split("/")[-1] full_path = os.path.join(save_dir, filename) if os.path.exists(full_path): os.remove(full_path) # 发送请求,自动处理压缩 resp = requests.get(url) # 获取正确编码,无法推断时默认UTF-8 encoding = resp.encoding if resp.encoding else "utf-8" # 按正确编码解码并保存 with open(full_path, "w", encoding=encoding) as f: f.write(resp.text) for url in sitemaps: download_and_save(url)
备选方案:保留wget并手动处理压缩
如果必须使用wget,可以先下载临时文件,判断是否为压缩文件后再解压保存:
import os import gzip import wget from django.conf import settings def download_and_save(url): save_dir = settings.STATICFILES_DIRS[0] filename = url.split("/")[-1] full_path = os.path.join(save_dir, filename) temp_path = f"{full_path}.tmp" if os.path.exists(full_path): os.remove(full_path) # 下载到临时文件 wget.download(url, temp_path) # 尝试解压,失败则直接重命名 try: with gzip.open(temp_path, "rb") as f_in: content = f_in.read() with open(full_path, "wb") as f_out: f_out.write(content) os.remove(temp_path) except OSError: os.rename(temp_path, full_path)
内容的提问来源于stack exchange,提问作者Michael Hawkins
相关产品推荐
相关产品推荐

