You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests下载维基百科图片失败的问题排查与解决求助

网页图片爬取问题排查与解决

我正在开发一个网页图片爬取项目,流程是先把图片链接写入文件,再用requests库批量下载到指定文件夹。最初用Google作为爬取源,后来换成维基百科。第一次尝试时大量图片打不开,改成按链接后缀命名后能打开的变多,但还是有不少打不开。奇怪的是,把这些打不开的链接拿到函数外单独下载完全正常,之后再用函数下载也能正常打开。常见图片后缀是svg.png和png。附上相关代码,求问题原因和解决方法。

批量下载函数代码

def download_images(file):
    object = file[0:file.index("IMAGELINKS") - 1]
    folder_name = object + "_images"
    dir = os.path.join("math_obj_images/original_images/", folder_name)
    if not os.path.exists(dir):
        os.mkdir(dir)
        with open("math_obj_image_links/" + file, "r") as f:
            count = 1
            for line in f:
                try:
                    if line[len(line) - 1] == "\n":
                        line = line[:len(line) - 1]
                    if line[0] != "/":
                        last_chunk = line.split("/")[len(line.split("/")) - 1]
                        endings = last_chunk.split(".")[1:]
                        image_ending = ""
                        for ending in endings:
                            image_ending += "." + ending
                        if image_ending == "":
                            continue
                        with open("math_obj_images/original_images/" + folder_name + "/" + object + str(count) + image_ending, "wb") as f:
                            f.write(requests.get(line).content)
                        file = object + "_IMAGEENDINGS.txt"
                        path = "math_obj_image_endings/" + file
                        with open(path, "a") as f:
                            f.write(image_ending + "\n")
                        count += 1
                except:
                    continue
            f.close()

单独下载可正常运行的代码

with open("test" + image_ending, "wb") as f:
    f.write(requests.get(line).content)

图片链接示例

  • https://upload.wikimedia.org/wikipedia/commons/thumb/6/63/Triangle.TrigArea.svg/120px-Triangle.TrigArea.svg.png
  • https://upload.wikimedia.org/wikipedia/commons/thumb/c/c9/Square_%28geometry%29.svg/120px-Square_%28geometry%29.svg.png
  • https://upload.wikimedia.org/wikipedia/commons/thumb/3/33/Hexahedron.png/120px-Hexahedron.png
  • https://upload.wikimedia.org/wikipedia/commons/thumb/2/22/Hypercube.svg/110px-Hypercube.svg.png
  • https://wikimedia.org/api/rest_v1/media/math/render/svg/5f8ab564115bf2f7f7d12a9f873d9c6c7a50190e
  • https://en.wikipedia.org/wiki/Special:CentralAutoLogin/start?type=1x1
  • https:/static/images/footer/wikimedia-button.png
  • https:/static/images/footer/poweredby_mediawiki_88x31.png

问题原因分析

  1. 反爬拦截:维基百科服务器会拦截无User-Agent的请求,返回非图片内容(如403页面),导致文件无法打开。单独请求时次数少或环境带默认UA未触发拦截,批量请求则触发反爬。
  2. 变量名冲突:函数内多次复用f作为文件句柄,导致句柄混乱,出现写入不完整的情况。
  3. 异常捕获过宽:except:捕获所有异常,掩盖了网络错误、文件写入错误等具体问题,无法针对性排查。
  4. 链接处理不严谨:存在格式错误的链接(如少一个/的https:/static/...)、非图片链接(如CentralAutoLogin的链接),下载内容自然不是图片。

解决方法

  1. 添加请求头:模拟浏览器请求,避免被反爬拦截,同时检查响应状态码确保请求成功:
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
response = requests.get(line, headers=headers)
if response.status_code == 200:
    f.write(response.content)
  1. 避免变量冲突:给不同的文件句柄用不同变量名,比如link_file、image_file、ending_file。
  2. 精准捕获异常:捕获具体异常类型并打印信息,方便排查:
try:
    # 核心逻辑
except requests.exceptions.RequestException as e:
    print(f"下载链接{line}失败:{e}")
    continue
except IOError as e:
    print(f"写入文件失败:{e}")
    continue
  1. 过滤无效链接:
    • 检查链接是否以http://或https://开头,过滤格式错误的链接;
    • 检查链接后缀或响应Content-Type头,过滤非图片链接;
  2. 优化路径拼接:用os.path.join拼接所有路径,避免手动拼接出错:
image_path = os.path.join(dir, f"{object}{count}{image_ending}")
with open(image_path, "wb") as image_file:
    image_file.write(response.content)

内容的提问来源于stack exchange,提问作者Preston Brown

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 08:06:24