You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

修复Python网页图片批量下载脚本的KeyError问题求助

问题描述

尝试编写Python函数从网页批量下载图片并重命名,但运行时触发KeyError: 'href',曾尝试将代码中的href改为data-profile,问题仍未解决。

原代码
import requests
from bs4 import BeautifulSoup
import os

#url = 'https://www.stocksy.com/service/contributors/'
def imagedown(url, folder):
    try:
        os.mkdir(os.path.join(os.getcwd(), folder))
    except:
        pass
    os.chdir(os.path.join(os.getcwd(), folder))
    r = requests.get(url)
    soup = BeautifulSoup(r.text, 'html.parser')
    images = soup.find_all('img')
    for image in images:
        name = image['href']
        link = image['src']
        with open(name.replace(' ', '-').replace('/', '') + '.jpg', 'wb') as f:
            im = requests.get(link)
            f.write(im.content)
            print('Writing: ', name)

imagedown('https://www.stocksy.com/service/contributors/', 'contributors')
报错信息
KeyError                                  Traceback (most recent call last)
Input In [4], in <cell line: 23>()
     20             f.write(im.content)
     21             print('Writing: ', name)
---> 23 imagedown('https://www.stocksy.com/service/contributors/', 'contributors')

Input In [4], in imagedown(url, folder)
     14 images = soup.find_all('img')
     15 for image in images:
---> 16     name = image['href']
     17     link = image['src']
     18     with open(name.replace(' ', '-').replace('/', '') + '.jpg', 'wb') as f:

File ~\anaconda3\lib\site-packages\bs4\element.py:1519, in Tag.__getitem__(self, key)
   1516 def __getitem__(self, key):
   1517     """tag[key] returns the value of the 'key' attribute for the Tag,
   1518     and throws an exception if it's not there."""
-> 1519     return self.attrs[key]

KeyError: 'href'
问题分析与修复
  1. 错误根源:img标签本身不存在href属性,你尝试替换的data-profile属于img的父级<a>标签属性,并非img自身属性,因此调用会报错。
  2. 图片命名优化:可以用img标签的alt属性作为文件名(如果存在),无alt则用索引兜底,避免命名异常。
  3. 图片链接处理:该网站图片采用懒加载机制,部分img的src是占位图,实际链接在data-src属性中;同时要补全相对链接为绝对链接,避免请求失败。
  4. 目录操作优化:避免直接切换工作目录,使用绝对路径操作文件更安全,同时用os.makedirs的exist_ok=True简化文件夹创建逻辑。
修正后的代码
import requests
from bs4 import BeautifulSoup
import os
from urllib.parse import urljoin

def imagedown(url, folder):
    # 创建目标文件夹,已存在则跳过
    folder_path = os.path.join(os.getcwd(), folder)
    os.makedirs(folder_path, exist_ok=True)
    
    # 请求网页,捕获请求失败情况
    r = requests.get(url)
    r.raise_for_status()
    soup = BeautifulSoup(r.text, 'html.parser')
    
    # 遍历所有img标签
    for idx, image in enumerate(soup.find_all('img')):
        # 获取图片实际链接:优先取data-src,无则取src
        img_link = image.get('data-src') or image.get('src')
        if not img_link:
            print(f"跳过无有效链接的图片")
            continue
        # 补全相对链接为绝对链接
        full_img_link = urljoin(url, img_link)
        
        # 生成合法文件名:优先用alt属性,无则用索引
        img_name = image.get('alt') or f"image_{idx}"
        # 清理文件名非法字符
        img_name = img_name.replace(' ', '-').replace('/', '').replace('\\', '') + '.jpg'
        img_save_path = os.path.join(folder_path, img_name)
        
        # 下载并保存图片
        try:
            img_response = requests.get(full_img_link)
            img_response.raise_for_status()
            with open(img_save_path, 'wb') as f:
                f.write(img_response.content)
            print(f"已保存:{img_name}")
        except Exception as e:
            print(f"下载失败 {full_img_link}:{str(e)}")

imagedown('https://www.stocksy.com/service/contributors/', 'contributors')

内容的提问来源于stack exchange,提问作者midomid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 11:45:34