You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python单URL图片爬虫代码实现多URL数组批量爬取

批量爬取多URL页面图片解决方案

原有代码问题

  • main函数中出现未定义的url22变量,循环逻辑错误
  • 没有适配多URL的处理流程
  • 文件夹创建每次需要手动输入,批量处理时操作繁琐
  • 文件夹创建失败后递归调用自身,存在死循环风险
  • 未处理页面内的相对路径图片链接,大概率出现爬取失败的情况

修改后可直接运行的完整代码

from bs4 import BeautifulSoup
import requests
import os

# 下载当前页面所有图片
def download_images(images, folder_name):
    count = 0
    print(f"当前页面共找到 {len(images)} 张图片")
    if len(images) == 0:
        print("当前页面无图片可下载")
        return
    for i, image in enumerate(images):
        image_link = None
        # 按优先级获取图片链接
        for attr in ["data-srcset", "data-src", "data-fallback-src", "src"]:
            if attr in image.attrs:
                image_link = image[attr]
                break
        if not image_link:
            continue
        # 补全相对路径图片链接
        if not image_link.startswith(("http://", "https://")):
            image_link = f"{current_url.rstrip('/')}/{image_link.lstrip('/')}"
        try:
            r = requests.get(image_link, timeout=10).content
            # 校验是否为图片资源
            try:
                str(r, 'utf-8')
            except UnicodeDecodeError:
                filename = image_link.strip().split('/')[-1].split('?')[0].strip()
                # 处理文件名非法字符,避免保存失败
                filename = "".join([c for c in filename if c not in r'\/:*?"<>|'])
                if not filename:
                    filename = f"image_{i+1}.jpg"
                save_path = os.path.join(folder_name, filename)
                with open(save_path, "wb+") as f:
                    f.write(r)
                count += 1
                print(f"已下载:{filename}")
        except Exception as e:
            print(f"下载图片失败 {image_link}: {str(e)}")
            continue
    print(f"当前页面下载完成:共 {count} 张下载成功,{len(images)-count} 张下载失败")

# 自动创建存储文件夹
def folder_create(index):
    folder_name = f"爬取结果_第{index}个页面"
    # 文件夹存在时自动加后缀,避免冲突
    suffix = 1
    while os.path.exists(folder_name):
        folder_name = f"爬取结果_第{index}个页面_{suffix}"
        suffix +=1
    os.mkdir(folder_name)
    return folder_name

def main(url_list):
    global current_url
    for index, url in enumerate(url_list, 1):
        print(f"\n===== 正在处理第{index}个URL:{url} =====")
        current_url = url
        try:
            r = requests.get(url, timeout=10)
            r.encoding = r.apparent_encoding
            soup = BeautifulSoup(r.text, 'html.parser')
            images = soup.find_all('img')
            folder_name = folder_create(index)
            download_images(images, folder_name)
        except Exception as e:
            print(f"访问URL失败 {url}: {str(e)}")
            continue

if __name__ == "__main__":
    input_urls = input("请输入需要爬取的URL,多个URL用英文逗号分隔:")
    url_list = [u.strip() for u in input_urls.split(',') if u.strip()]
    if not url_list:
        print("未输入有效URL")
    else:
        main(url_list)

功能说明

  • 支持输入多个URL,用英文逗号分隔即可
  • 自动为每个URL创建独立的存储文件夹,避免图片重名覆盖
  • 自动补全页面中的相对路径图片链接,提升爬取成功率
  • 自动处理非法文件名,避免保存失败
  • 增加超时设置,避免长时间无响应卡住

使用方法

运行代码后按提示输入多个URL,格式示例:https://example.com,https://example2.com,等待爬取完成后,每个页面的图片会保存在对应序号的「爬取结果」文件夹中。

内容的提问来源于stack exchange,提问作者Dragonasu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 13:54:04