使用Python Beautiful Soup爬取网站图片报错,请求问题排查
台湾高铁时刻表页面图片爬取报错解决
问题描述
尝试爬取台湾高铁时刻表页面(https://www.thsrc.com.tw/tw/TimeTable/SearchResult)的图片时,运行提供的Python代码出现报错。
报错原因分析
结合代码逻辑和网站特性,常见报错原因包括:
- 反爬拦截:未携带浏览器请求头,被网站识别为爬虫,导致返回页面不完整或403禁止访问。
- 保存目录缺失:代码直接写入
save_image/目录,但该目录未提前创建,触发文件路径错误。 - URL拼接错误:部分图片
src可能已是完整URL或为空值,直接拼接域名会生成无效地址。 - 下载工具兼容性:
urlretrieve无法复用requests的会话和请求头,容易被反爬机制拦截。
解决方案
针对上述问题,采取以下修正措施:
- 添加浏览器请求头:模拟真实用户访问,绕过基础反爬。
- 自动创建保存目录:提前检查并创建目标文件夹,避免路径错误。
- 优化URL处理逻辑:判断
src是否为完整URL,过滤空值,确保地址有效。 - 改用
requests下载图片:复用会话请求头,提升下载成功率。
修正后的代码
import requests from bs4 import BeautifulSoup import os # 自动创建保存目录 save_directory = 'save_image' if not os.path.exists(save_directory): os.makedirs(save_directory) # 模拟浏览器请求头 request_headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } # 发起页面请求 target_url = 'https://www.thsrc.com.tw/tw/TimeTable/SearchResult' page_response = requests.get(target_url, headers=request_headers) page_response.encoding = 'utf-8' # 解析页面内容 soup = BeautifulSoup(page_response.text, 'html.parser') image_elements = soup.find_all('img') # 遍历并下载图片 for idx, img in enumerate(image_elements): # 跳过第一个图片(可根据需求调整) if idx == 0: continue img_src = img.get('src') if not img_src: print('跳过无src属性的图片') continue # 处理图片URL if img_src.startswith(('http://', 'https://')): img_url = img_src else: img_url = f'https://www.thsrc.com.tw{img_src}' img_filename = img_src.split('/')[-1] save_path = os.path.join(save_directory, img_filename) try: # 下载图片 img_response = requests.get(img_url, headers=request_headers, stream=True) img_response.raise_for_status() with open(save_path, 'wb') as img_file: for chunk in img_response.iter_content(chunk_size=1024): img_file.write(chunk) print(f'成功保存:{img_filename}') except Exception as err: print(f'下载失败 {img_filename}:{str(err)}')
内容的提问来源于stack exchange,提问作者Tinny
相关产品推荐
相关产品推荐

