如何使用Python的Selenium提取元素style属性中的URL链接
HTML背景图URL提取方案
以下是不同场景下的实现方式:
方法1:Python 环境(推荐,解析HTML结构更稳定)
依赖BeautifulSoup库解析DOM结构,避免正则匹配的容错性问题:
from bs4 import BeautifulSoup import re # 填入你的HTML片段 html_raw = ''' <picture data-testid="menu-product-image" style="background-image: url("https://micro-assets.foodora.com/img/logo-simple-fp.svg");"><div class="photo" style="background-image: url("https://images.deliveryhero.io/image/fd-bd/Products/1032643.jpg?width=200");"></div></picture> ''' # 解析HTML soup = BeautifulSoup(html_raw, 'html.parser') # 定位目标div元素 target_div = soup.find('picture', {'data-testid': 'menu-product-image'}).find('div', class_='photo') # 提取style属性中的URL style_content = target_div['style'] target_url = re.search(r'url\("?([^")]+)"?\)', style_content).group(1) print(target_url) # 输出结果:https://images.deliveryhero.io/image/fd-bd/Products/1032643.jpg?width=200
方法2:浏览器前端JavaScript环境
直接用原生DOM选择器定位元素提取:
// 定位元素 const photoDiv = document.querySelector('picture[data-testid="menu-product-image"] .photo') // 提取背景图URL const bgStyle = photoDiv.style.backgroundImage const targetUrl = bgStyle.match(/url\("?([^")]+)"?\)/)[1] console.log(targetUrl)
方法3:纯正则匹配(仅适用于固定格式的文本片段)
如果仅需要对固定格式的文本做快速提取,可以用以下正则匹配,捕获组1即为目标URL:
url\("(https://images\.deliveryhero\.io[^"]+)"\)
注意:纯正则匹配不兼容HTML结构变动的场景,复杂HTML提取优先选择DOM解析方案。
内容的提问来源于stack exchange,提问作者Samyak jain
相关产品推荐
相关产品推荐

