You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从HTML页面提取内嵌JSON格式的门店位置数据?

解决丝芙兰门店页面JSON解析失败问题

我尝试从丝芙兰门店列表页面获取特定州的门店位置,但请求返回的内容类型为text/html,直接调用response.json()时触发了JSONDecodeError异常。排查后发现页面的JSON数据内嵌在type="text/json"的script标签中,通过解析HTML提取该标签内容再转换为JSON即可解决问题。

初始请求代码

import requests
import json
headers = {
    'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36',
    'accept-language': 'en-US,en;q=0.9',
    'Content-type': 'application/json', 
    'Accept': 'text/plain'
}

response = requests.get('https://www.sephora.com/happening/storelist', headers=headers)

报错信息

requests.exceptions.JSONDecodeError: Expecting value: line 1 column 1 (char 0)

解决代码

先确保安装beautifulsoup4库:pip install beautifulsoup4,然后用以下代码提取JSON数据:

from bs4 import BeautifulSoup

soup = BeautifulSoup(response.text, 'html.parser')
script_tag = soup.find('script', {'type': 'text/json'})  
if script_tag:
    specific_content = script_tag.text
    json_data = json.loads(specific_content)
    # 后续可根据需求从json_data中筛选特定州的门店信息
else:
    print("未找到目标script标签。")

内容的提问来源于stack exchange,提问作者mhj.DA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 12:12:39