You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python谷歌图片爬取脚本无报错终止问题求助

问题分析与解决方案

你的脚本在调用download_image函数后无响应终止,核心原因是:

  • requests.get()未设置超时,当目标服务器无响应时,请求会无限挂起
  • 未设置合法的User-Agent请求头,部分服务器会拒绝无标识的请求,甚至触发反爬策略导致请求被阻塞
  • 连续高频请求容易触发目标网站的反爬限制,导致请求被挂起

修复后的代码

1. 优化download_image函数

import requests
import os
import time

def download_image(file_path, url, file_name):
    # 设置模拟浏览器的请求头
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    try:
        # 添加超时参数,避免请求无限等待
        response = requests.get(url, headers=headers, timeout=(10, 30))
        response.raise_for_status()
        
        # 确保保存图片的文件夹存在
        os.makedirs(file_path, exist_ok=True)
        
        with open(os.path.join(file_path, file_name), 'wb') as file:
            file.write(response.content)
        print(f"Image downloaded successfully to {os.path.join(file_path, file_name)}")
    except requests.exceptions.HTTPError as http_error:
        print(f"HTTP error occurred: {http_error}")
    except requests.exceptions.Timeout:
        print(f"Request timed out for URL: {url}")
    except Exception as error:
        print(f"An error occurred: {error}")
    # 添加请求间隔,降低反爬风险
    time.sleep(1)

2. 优化调用逻辑

def enhanced_dataset_folder(name:str, tag:str, df):
    DRIVER_PATH = "chromedriver"
    wd = webdriver.Chrome(DRIVER_PATH)
    urls = get_images(tag, wd, 1, 2)
    folder_name = name.split('/')[0]
    props = tag.split(' ')
    
    for i, url in enumerate(urls):
        try:
            img_name = f"{i}_img{i}.jpg"
            download_image(f"train/{folder_name}/", url, img_name)
        except Exception as e:
            print('Fail: ', e)
            continue
        else:
            print("ok")
            # df.append([f"{folder_name}/{img_name}", tag, props[0], props[1], props[2]], ignore_index=True)
    wd.quit()

关键优化点说明

  • 超时设置:给requests.get添加timeout参数,确保请求在指定时间内结束,避免脚本无限挂起
  • 请求头模拟:使用浏览器标识的User-Agent,避免被服务器识别为爬虫而拒绝服务
  • 文件夹预创建:自动创建目标文件夹,避免因路径不存在导致的写入失败
  • 请求间隔:添加1秒等待,降低请求频率,减少触发反爬策略的概率
  • 细分异常:单独捕获超时异常,更精准定位问题

内容的提问来源于stack exchange,提问作者N7Legend

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 09:45:10