You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup抓取嵌套Div中Imgur图库图片失败求助

问题分析与解决

你当前代码的核心问题是请求的目标是一张图片文件(https://i.sstatic.net/mHf6v.jpg),而非Imgur图库的HTML网页。用BeautifulSoup解析图片二进制内容,自然找不到任何HTML元素(比如div.PostContent-imageWrapper-rounded),所以会输出"Div not found"。

修正步骤

1. 替换为正确的Imgur图库页面URL

你提供的链接是Stack Exchange上的截图,需要换成你个人Imgur图库的实际网页地址(比如类似https://imgur.com/user/你的用户名/posts这样的格式)。

2. 修正代码逻辑(静态页面场景)

以下是调整后的代码,针对静态加载的Imgur图库页面:

import requests
from bs4 import BeautifulSoup

# 替换为你的Imgur图库页面URL
target_url = "https://imgur.com/user/你的用户名/posts"
r = requests.get(target_url)
r.encoding = 'utf-8'

soup = BeautifulSoup(r.text, 'html.parser')

# 查找所有图片容器(需根据Imgur实际页面结构调整选择器)
img_containers = soup.find_all("div", class_="post")

for container in img_containers:
    img_tag = container.find("img", class_="post-image")
    if img_tag:
        img_src = img_tag.get('src')
        # 补全协议头(如果链接是相对路径)
        if img_src.startswith("//"):
            img_src = "https:" + img_src
        print(img_src)

3. 处理动态加载场景

如果Imgur图库使用JavaScript滚动加载图片,requests只能获取初始页面内容,这时候需要用selenium模拟浏览器渲染:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("https://imgur.com/user/你的用户名/posts")

# 等待初始图片加载完成
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CLASS_NAME, "post-image")))

# 获取所有可见图片链接
img_tags = driver.find_elements(By.CLASS_NAME, "post-image")
for img in img_tags:
    print(img.get_attribute("src"))

driver.quit()

关键注意事项

  • 必须确保请求的是HTML网页,而非图片文件;
  • Imgur页面结构可能更新,需用浏览器开发者工具(F12)查看最新的元素类名或选择器;
  • 频繁请求可能触发反爬机制,建议添加User-Agent请求头并控制请求频率。

内容的提问来源于stack exchange,提问作者Donchev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 23:05:17