如何用Python的requests与BeautifulSoup提取网页图片URL
提取背景图片URL的解决方案
你可以通过正则表达式或者字符串分割的方式,从style属性中提取出图片URL,以下是具体实现:
方法一:正则表达式(推荐,适配更复杂的格式)
利用正则匹配url(...)结构中的内容,代码修改如下:
import requests from bs4 import BeautifulSoup import re url = "https://www.mayde.com/ponytail" res = requests.get(url) res.raise_for_status() soup = BeautifulSoup(res.text, "lxml") item_address1 = soup.find_all("div", attrs={"class":"BEmRci heightByImageRatio heightByImageRatio4"}) # 提取style文本 style_content = item_address1[0]["style"] # 匹配url(...)中的URL部分 img_url_match = re.search(r'url\((.*?)\)', style_content) if img_url_match: pure_img_url = img_url_match.group(1) print(pure_img_url)
正则表达式r'url\((.*?)\)'的作用:
url\(:精准匹配url(字符串(括号前加转义符\因为括号是正则特殊字符)(.*?):非贪婪模式匹配括号内的所有内容,直到遇到第一个)\):匹配闭合的)
方法二:字符串分割(简单直接,适合固定格式)
如果确定style格式固定,也可以用字符串分割快速提取:
# 接原代码获取style_content之后 pure_img_url = style_content.split('url(')[1].split(')')[0] print(pure_img_url)
这种方式通过两次分割,先以url(拆分取后半部分,再以)拆分取前半部分,直接得到目标URL。
内容的提问来源于stack exchange,提问作者Peter Yun
相关产品推荐
相关产品推荐

