Python如何获取div标签的style属性值并提取其中的图片链接
解决方案
方案1:BeautifulSoup提取(适用于静态页面场景)
该方案适用于目标div标签直接存在于页面静态源码中的情况,操作步骤如下:
- 用class属性作为筛选条件定位div,不需要关联标签文本内容
- 获取元素的style属性后用正则匹配提取图片链接
示例代码:
from bs4 import BeautifulSoup import re # html_content为你爬取到的页面源码 soup = BeautifulSoup(html_content, 'lxml') # 按指定class定位目标div target_div = soup.find('div', class_='v-image__image v-image__image--cover') if target_div: # 提取style属性值 style_text = target_div.get('style') # 正则匹配url中的图片链接 img_url = re.search(r'url\("(.*?)"\)', style_text).group(1) print(img_url)
如果soup.find返回空,可先打印完整页面源码确认该div是否存在于静态内容中,不存在则说明是动态渲染页面,需使用下方Selenium方案。
方案2:Chrome Driver(Selenium)提取(适用于动态渲染页面场景)
该方案适用于页面内容由JS动态生成的场景,需等待元素加载完成后再操作:
- 添加显式等待逻辑,确保目标元素已经加载到DOM树中
- 定位到元素后直接获取style属性,再正则提取链接
示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import re driver = webdriver.Chrome() driver.get("替换为你的目标页面URL") # 显式等待最多10秒,直到目标div加载完成 wait = WebDriverWait(driver, 10) target_div = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div.v-image__image.v-image__image--cover"))) # 获取style属性值 style_text = target_div.get_attribute("style") # 正则提取图片链接 img_url = re.search(r'url\("(.*?)"\)', style_text).group(1) print(img_url) driver.quit()
常见问题排查
- 元素属性的获取和标签首尾是否有文本没有任何关联,只要元素存在于DOM树中就可以正常读取属性
- 如果定位失败,先检查页面是否嵌套iframe,如有需要先调用
driver.switch_to.frame()切换到对应iframe后再定位 - 若实际style中url使用单引号/无引号包裹,调整正则匹配规则即可
内容的提问来源于stack exchange,提问作者pickle rick
相关产品推荐
相关产品推荐

