使用BeautifulSoup抓取嵌套Div中Imgur图库图片失败求助
问题分析与解决
你当前代码的核心问题是请求的目标是一张图片文件(https://i.sstatic.net/mHf6v.jpg),而非Imgur图库的HTML网页。用BeautifulSoup解析图片二进制内容,自然找不到任何HTML元素(比如div.PostContent-imageWrapper-rounded),所以会输出"Div not found"。
修正步骤
1. 替换为正确的Imgur图库页面URL
你提供的链接是Stack Exchange上的截图,需要换成你个人Imgur图库的实际网页地址(比如类似https://imgur.com/user/你的用户名/posts这样的格式)。
2. 修正代码逻辑(静态页面场景)
以下是调整后的代码,针对静态加载的Imgur图库页面:
import requests from bs4 import BeautifulSoup # 替换为你的Imgur图库页面URL target_url = "https://imgur.com/user/你的用户名/posts" r = requests.get(target_url) r.encoding = 'utf-8' soup = BeautifulSoup(r.text, 'html.parser') # 查找所有图片容器(需根据Imgur实际页面结构调整选择器) img_containers = soup.find_all("div", class_="post") for container in img_containers: img_tag = container.find("img", class_="post-image") if img_tag: img_src = img_tag.get('src') # 补全协议头(如果链接是相对路径) if img_src.startswith("//"): img_src = "https:" + img_src print(img_src)
3. 处理动态加载场景
如果Imgur图库使用JavaScript滚动加载图片,requests只能获取初始页面内容,这时候需要用selenium模拟浏览器渲染:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("https://imgur.com/user/你的用户名/posts") # 等待初始图片加载完成 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "post-image"))) # 获取所有可见图片链接 img_tags = driver.find_elements(By.CLASS_NAME, "post-image") for img in img_tags: print(img.get_attribute("src")) driver.quit()
关键注意事项
- 必须确保请求的是HTML网页,而非图片文件;
- Imgur页面结构可能更新,需用浏览器开发者工具(F12)查看最新的元素类名或选择器;
- 频繁请求可能触发反爬机制,建议添加
User-Agent请求头并控制请求频率。
内容的提问来源于stack exchange,提问作者Donchev
相关产品推荐
相关产品推荐

