You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取图片时触发KeyError: 'src'错误求助

解决爬取漫画图片时的KeyError: 'src'问题

问题描述

尝试使用Python的requests和BeautifulSoup库爬取指定漫画网站的图片,代码如下:

import requests
from bs4 import BeautifulSoup
import os


URL = "https://ww0.jujutsukaisen.online/manga/jujutsu-kaisen-manga-chapter-1/?2022-12-29"

r = requests.get(URL)
soup = BeautifulSoup(r.text, 'html.parser')

images = soup.find_all('img')

for image in images:
    name='name'
    link = image['src']
    print(link)

运行时出现错误:

C:\Users\amicr\AppData\Local\Programs\Python\Python310\python.exe D:\Junk\Scrape.py 
Traceback (most recent call last):
  File "D:\Junk\Scrape.py", line 17, in <module>
    link = image['src']
  File "C:\Users\amicr\AppData\Local\Programs\Python\Python310\lib\site-packages\bs4\element.py", line 1519, in __getitem__
    return self.attrs[key]
KeyError: 'src'

错误原因

  1. 页面内并非所有<img>标签都包含src属性,部分标签是占位元素,或是采用懒加载机制(实际图片链接存储在data-src等其他属性中),直接通过image['src']强制访问不存在的属性会触发KeyError。
  2. 目标网站的漫画图片链接可能并未直接放在src属性里,需要查看网页实际元素结构确认存储位置。

修复方案

1. 安全获取图片链接并下载

使用get()方法安全获取属性,同时兼容常见的懒加载属性,补充下载逻辑并处理相对路径:

import requests
from bs4 import BeautifulSoup
import os
from urllib.parse import urljoin

URL = "https://ww0.jujutsukaisen.online/manga/jujutsu-kaisen-manga-chapter-1/?2022-12-29"
base_url = "https://ww0.jujutsukaisen.online"

# 创建图片保存目录
save_dir = "jujutsu_chapter_1"
if not os.path.exists(save_dir):
    os.makedirs(save_dir)

# 添加请求头模拟浏览器访问,避免反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
r = requests.get(URL, headers=headers)
soup = BeautifulSoup(r.text, 'html.parser')

images = soup.find_all('img')

for idx, image in enumerate(images):
    # 优先取src,无则尝试常见的懒加载属性
    img_link = image.get('src') or image.get('data-src') or image.get('data-url')
    if not img_link:
        continue  # 跳过无有效链接的img标签
    
    # 将相对路径转为绝对路径
    img_link = urljoin(base_url, img_link)
    
    # 下载并保存图片
    try:
        img_data = requests.get(img_link, headers=headers).content
        img_name = f"page_{idx+1}.jpg"
        with open(os.path.join(save_dir, img_name), 'wb') as f:
            f.write(img_data)
        print(f"已下载: {img_name}")
    except Exception as e:
        print(f"下载失败 {img_link}: {str(e)}")

2. 精准定位漫画图片

很多漫画网站的图片会放在特定类名的容器中,可以缩小查找范围,避免无关的img标签:

# 替换原有的images查找代码,改为定位特定容器内的img(class名需根据实际网页调整)
images = soup.find_all('img', class_='chapter-img')

注意事项

  • 爬取网站内容前,请遵守网站的robots.txt协议及相关法律法规,本次操作仅用于学习目的。
  • 若仍无法获取图片,需检查网页是否通过JavaScript动态加载内容,这种情况下可能需要使用Selenium等工具渲染页面。

内容的提问来源于stack exchange,提问作者ADITYA GHOSH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 20:55:48