修复Python网页图片批量下载脚本的KeyError问题求助
问题描述
尝试编写Python函数从网页批量下载图片并重命名,但运行时触发KeyError: 'href',曾尝试将代码中的href改为data-profile,问题仍未解决。
原代码
import requests from bs4 import BeautifulSoup import os #url = 'https://www.stocksy.com/service/contributors/' def imagedown(url, folder): try: os.mkdir(os.path.join(os.getcwd(), folder)) except: pass os.chdir(os.path.join(os.getcwd(), folder)) r = requests.get(url) soup = BeautifulSoup(r.text, 'html.parser') images = soup.find_all('img') for image in images: name = image['href'] link = image['src'] with open(name.replace(' ', '-').replace('/', '') + '.jpg', 'wb') as f: im = requests.get(link) f.write(im.content) print('Writing: ', name) imagedown('https://www.stocksy.com/service/contributors/', 'contributors')
报错信息
KeyError Traceback (most recent call last) Input In [4], in <cell line: 23>() 20 f.write(im.content) 21 print('Writing: ', name) ---> 23 imagedown('https://www.stocksy.com/service/contributors/', 'contributors') Input In [4], in imagedown(url, folder) 14 images = soup.find_all('img') 15 for image in images: ---> 16 name = image['href'] 17 link = image['src'] 18 with open(name.replace(' ', '-').replace('/', '') + '.jpg', 'wb') as f: File ~\anaconda3\lib\site-packages\bs4\element.py:1519, in Tag.__getitem__(self, key) 1516 def __getitem__(self, key): 1517 """tag[key] returns the value of the 'key' attribute for the Tag, 1518 and throws an exception if it's not there.""" -> 1519 return self.attrs[key] KeyError: 'href'
问题分析与修复
- 错误根源:
img标签本身不存在href属性,你尝试替换的data-profile属于img的父级<a>标签属性,并非img自身属性,因此调用会报错。 - 图片命名优化:可以用
img标签的alt属性作为文件名(如果存在),无alt则用索引兜底,避免命名异常。 - 图片链接处理:该网站图片采用懒加载机制,部分
img的src是占位图,实际链接在data-src属性中;同时要补全相对链接为绝对链接,避免请求失败。 - 目录操作优化:避免直接切换工作目录,使用绝对路径操作文件更安全,同时用
os.makedirs的exist_ok=True简化文件夹创建逻辑。
修正后的代码
import requests from bs4 import BeautifulSoup import os from urllib.parse import urljoin def imagedown(url, folder): # 创建目标文件夹,已存在则跳过 folder_path = os.path.join(os.getcwd(), folder) os.makedirs(folder_path, exist_ok=True) # 请求网页,捕获请求失败情况 r = requests.get(url) r.raise_for_status() soup = BeautifulSoup(r.text, 'html.parser') # 遍历所有img标签 for idx, image in enumerate(soup.find_all('img')): # 获取图片实际链接:优先取data-src,无则取src img_link = image.get('data-src') or image.get('src') if not img_link: print(f"跳过无有效链接的图片") continue # 补全相对链接为绝对链接 full_img_link = urljoin(url, img_link) # 生成合法文件名:优先用alt属性,无则用索引 img_name = image.get('alt') or f"image_{idx}" # 清理文件名非法字符 img_name = img_name.replace(' ', '-').replace('/', '').replace('\\', '') + '.jpg' img_save_path = os.path.join(folder_path, img_name) # 下载并保存图片 try: img_response = requests.get(full_img_link) img_response.raise_for_status() with open(img_save_path, 'wb') as f: f.write(img_response.content) print(f"已保存:{img_name}") except Exception as e: print(f"下载失败 {full_img_link}:{str(e)}") imagedown('https://www.stocksy.com/service/contributors/', 'contributors')
内容的提问来源于stack exchange,提问作者midomid
相关产品推荐
相关产品推荐

