如何使用BeautifulSoup提取网页gtag/jQuery代码内的price价格值
提取price字段的实现方案
因为你已确认页面内price是唯一命名字段,直接使用「BeautifulSoup提取script标签内容+正则匹配」即可完成提取,无需引入额外依赖。
完整修改后代码
import requests import re from bs4 import BeautifulSoup URL = '' # 替换为你的目标页面地址 headers = { "User-Agent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36 Edg/93.0.961.44'} page = requests.get(URL, headers=headers) soup = BeautifulSoup(page.content, 'html.parser') # 提取所有script标签的文本内容 script_content = '' for script in soup.find_all('script'): if script.string: script_content += script.string # 正则匹配唯一的price值 price_match = re.search(r"'price'\s*:\s*'(\d+\.?\d*)'", script_content) if price_match: price = price_match.group(1) # 如需数值类型直接转float即可 # price = float(price_match.group(1)) print(f"提取到的价格为:{price}") else: print("未匹配到price字段")
说明
- 正则中的
\s*是为了兼容冒号前后可能存在的任意空格、换行符,避免格式变化导致匹配失败 - 如果后续页面
price字段不再唯一,可以调整正则的匹配范围,先定位到gtag事件的代码块后再提取价格
内容的提问来源于stack exchange,提问作者Nick Winston
相关产品推荐
相关产品推荐

