用BeautifulSoup抓取网页时,如何获取CSS::before后的img元素src?
解决BeautifulSoup获取房产条目最后一张img的src问题
你遇到的核心问题是无法获取房产条目里除第一张外的图片src,这里需要明确:CSS ::before 是伪元素,不会出现在实际DOM结构中,BeautifulSoup无法抓取伪元素,但这不影响对真实img元素的获取,问题大概率是选择器使用或页面动态加载导致的。
下面给出两种解决方案:
方案一:静态HTML抓取(适用于图片在初始HTML中)
修改你的循环逻辑,直接定位每个房产条目下的所有img元素,取最后一个的src属性:
import requests from bs4 import BeautifulSoup s = requests.session() r = s.get("https://www.immobiliare.it/vendita-case/milano/forlanini/?criterio=dataModifica&ordine=desc") soup = BeautifulSoup(r.content, "lxml") # 避免变量名与内置关键字冲突,改用property_item for property_item in soup.find_all("li", {"class": "nd-list__item in-realEstateResults__item"}): # 获取当前条目下所有img元素 all_imgs = property_item.find_all("img") if all_imgs: # 取最后一张图片的src属性,用get()避免无属性时报错 last_img_src = all_imgs[-1].get("src") print(last_img_src) else: print("该房产条目无图片")
方案二:动态页面抓取(适用于图片通过JS动态加载)
如果初始HTML里没有后续图片,说明图片是JavaScript动态渲染的,这时候需要用Selenium模拟浏览器加载完整页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量) driver = webdriver.Chrome() driver.get("https://www.immobiliare.it/vendita-case/milano/forlanini/?criterio=dataModifica&ordine=desc") # 等待房产条目加载完成,最长等待10秒 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "li.nd-list__item.in-realEstateResults__item")) ) # 获取渲染后的完整HTML soup = BeautifulSoup(driver.page_source, "lxml") for property_item in soup.find_all("li", {"class": "nd-list__item in-realEstateResults__item"}): all_imgs = property_item.find_all("img") if all_imgs: last_img_src = all_imgs[-1].get("src") print(last_img_src) else: print("该房产条目无图片") # 关闭浏览器 driver.quit()
关键说明
- 不要被
::before伪元素干扰,它只是页面样式的一部分,不会遮挡真实的img元素 - 使用
get("src")替代直接["src"],可以避免因部分图片无src属性导致的报错 - 如果静态抓取无效,优先考虑页面动态加载的情况,改用Selenium方案
内容的提问来源于stack exchange,提问作者Kaynef21
相关产品推荐
相关产品推荐

