You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python wget下载XML正常,但TXT/HTML文件编码异常无法读取

问题分析与解决方案

核心原因推测

你遇到的问题大概率是以下两种情况之一:

  1. TXT/HTML文件被服务器压缩传输:服务器对TXT、HTML返回了gzip/deflate压缩后的二进制内容,但wget或urllib未自动解压,直接保存的压缩文件自然会显示乱码;而XML文件可能未被服务器压缩,因此正常。
  2. 编码不匹配:服务器返回的TXT/HTML使用了非UTF-8编码(如GBK、GB2312、UTF-16等),但你强制用UTF-8读取,导致编码解析错误。

验证方法

先通过简单代码排查问题根源:

import requests
# 替换成你的TXT/HTML链接
url = "https://.../urllist.txt"
resp = requests.get(url)
# 打印响应头,重点看Content-Encoding和Content-Type
print("响应头:", resp.headers)
# 打印自动推断的编码
print("推断编码:", resp.encoding)
  • 如果Content-Encoding显示gzip或deflate,说明是压缩问题;
  • 如果Content-Type里的charset不是utf-8,说明是编码不匹配问题。

解决方案:使用Requests库处理(推荐)

requests库会自动处理gzip/deflate压缩,并能从响应头或内容中自动推断正确编码,完美解决这两类问题。修改后的代码如下:

import os
import requests
from django.conf import settings

sitemaps = [
    "https://.../sitemap.xml",
    "https://.../sitemap_images.xml",
    "https://.../sitemap_video.xml",
    "https://.../sitemap_mobile.xml",
    "https://.../sitemap.html",
    "https://.../urllist.txt",
    "https://.../ror.xml"
]

def download_and_save(url):
    save_dir = settings.STATICFILES_DIRS[0]
    filename = url.split("/")[-1]
    full_path = os.path.join(save_dir, filename)
    
    if os.path.exists(full_path):
        os.remove(full_path)
    
    # 发送请求,自动处理压缩
    resp = requests.get(url)
    # 获取正确编码,无法推断时默认UTF-8
    encoding = resp.encoding if resp.encoding else "utf-8"
    
    # 按正确编码解码并保存
    with open(full_path, "w", encoding=encoding) as f:
        f.write(resp.text)

for url in sitemaps:
    download_and_save(url)

备选方案:保留wget并手动处理压缩

如果必须使用wget,可以先下载临时文件,判断是否为压缩文件后再解压保存:

import os
import gzip
import wget
from django.conf import settings

def download_and_save(url):
    save_dir = settings.STATICFILES_DIRS[0]
    filename = url.split("/")[-1]
    full_path = os.path.join(save_dir, filename)
    temp_path = f"{full_path}.tmp"
    
    if os.path.exists(full_path):
        os.remove(full_path)
    
    # 下载到临时文件
    wget.download(url, temp_path)
    
    # 尝试解压,失败则直接重命名
    try:
        with gzip.open(temp_path, "rb") as f_in:
            content = f_in.read()
        with open(full_path, "wb") as f_out:
            f_out.write(content)
        os.remove(temp_path)
    except OSError:
        os.rename(temp_path, full_path)

内容的提问来源于stack exchange,提问作者Michael Hawkins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 04:05:24