You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Beautiful Soup爬取网站图片报错,请求问题排查

台湾高铁时刻表页面图片爬取报错解决

问题描述

尝试爬取台湾高铁时刻表页面(https://www.thsrc.com.tw/tw/TimeTable/SearchResult)的图片时,运行提供的Python代码出现报错。

报错原因分析

结合代码逻辑和网站特性,常见报错原因包括:

  • 反爬拦截:未携带浏览器请求头,被网站识别为爬虫,导致返回页面不完整或403禁止访问。
  • 保存目录缺失:代码直接写入save_image/目录,但该目录未提前创建,触发文件路径错误。
  • URL拼接错误:部分图片src可能已是完整URL或为空值,直接拼接域名会生成无效地址。
  • 下载工具兼容性:urlretrieve无法复用requests的会话和请求头,容易被反爬机制拦截。

解决方案

针对上述问题,采取以下修正措施:

  1. 添加浏览器请求头:模拟真实用户访问,绕过基础反爬。
  2. 自动创建保存目录:提前检查并创建目标文件夹,避免路径错误。
  3. 优化URL处理逻辑:判断src是否为完整URL,过滤空值,确保地址有效。
  4. 改用requests下载图片:复用会话请求头,提升下载成功率。

修正后的代码

import requests
from bs4 import BeautifulSoup
import os

# 自动创建保存目录
save_directory = 'save_image'
if not os.path.exists(save_directory):
    os.makedirs(save_directory)

# 模拟浏览器请求头
request_headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}

# 发起页面请求
target_url = 'https://www.thsrc.com.tw/tw/TimeTable/SearchResult'
page_response = requests.get(target_url, headers=request_headers)
page_response.encoding = 'utf-8'

# 解析页面内容
soup = BeautifulSoup(page_response.text, 'html.parser')
image_elements = soup.find_all('img')

# 遍历并下载图片
for idx, img in enumerate(image_elements):
    # 跳过第一个图片(可根据需求调整)
    if idx == 0:
        continue
    
    img_src = img.get('src')
    if not img_src:
        print('跳过无src属性的图片')
        continue
    
    # 处理图片URL
    if img_src.startswith(('http://', 'https://')):
        img_url = img_src
    else:
        img_url = f'https://www.thsrc.com.tw{img_src}'
    
    img_filename = img_src.split('/')[-1]
    save_path = os.path.join(save_directory, img_filename)
    
    try:
        # 下载图片
        img_response = requests.get(img_url, headers=request_headers, stream=True)
        img_response.raise_for_status()
        
        with open(save_path, 'wb') as img_file:
            for chunk in img_response.iter_content(chunk_size=1024):
                img_file.write(chunk)
        print(f'成功保存:{img_filename}')
    except Exception as err:
        print(f'下载失败 {img_filename}:{str(err)}')

内容的提问来源于stack exchange,提问作者Tinny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 09:54:14