Python谷歌图片爬取脚本无报错终止问题求助
问题分析与解决方案
你的脚本在调用download_image函数后无响应终止,核心原因是:
requests.get()未设置超时,当目标服务器无响应时,请求会无限挂起- 未设置合法的
User-Agent请求头,部分服务器会拒绝无标识的请求,甚至触发反爬策略导致请求被阻塞 - 连续高频请求容易触发目标网站的反爬限制,导致请求被挂起
修复后的代码
1. 优化download_image函数
import requests import os import time def download_image(file_path, url, file_name): # 设置模拟浏览器的请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: # 添加超时参数,避免请求无限等待 response = requests.get(url, headers=headers, timeout=(10, 30)) response.raise_for_status() # 确保保存图片的文件夹存在 os.makedirs(file_path, exist_ok=True) with open(os.path.join(file_path, file_name), 'wb') as file: file.write(response.content) print(f"Image downloaded successfully to {os.path.join(file_path, file_name)}") except requests.exceptions.HTTPError as http_error: print(f"HTTP error occurred: {http_error}") except requests.exceptions.Timeout: print(f"Request timed out for URL: {url}") except Exception as error: print(f"An error occurred: {error}") # 添加请求间隔,降低反爬风险 time.sleep(1)
2. 优化调用逻辑
def enhanced_dataset_folder(name:str, tag:str, df): DRIVER_PATH = "chromedriver" wd = webdriver.Chrome(DRIVER_PATH) urls = get_images(tag, wd, 1, 2) folder_name = name.split('/')[0] props = tag.split(' ') for i, url in enumerate(urls): try: img_name = f"{i}_img{i}.jpg" download_image(f"train/{folder_name}/", url, img_name) except Exception as e: print('Fail: ', e) continue else: print("ok") # df.append([f"{folder_name}/{img_name}", tag, props[0], props[1], props[2]], ignore_index=True) wd.quit()
关键优化点说明
- 超时设置:给
requests.get添加timeout参数,确保请求在指定时间内结束,避免脚本无限挂起 - 请求头模拟:使用浏览器标识的
User-Agent,避免被服务器识别为爬虫而拒绝服务 - 文件夹预创建:自动创建目标文件夹,避免因路径不存在导致的写入失败
- 请求间隔:添加1秒等待,降低请求频率,减少触发反爬策略的概率
- 细分异常:单独捕获超时异常,更精准定位问题
内容的提问来源于stack exchange,提问作者N7Legend
相关产品推荐
相关产品推荐

