如何用Python的json.loads解析HTML中的window.__PRELOADED_STATE__数据
提取HTML中预加载状态的解决方案
- 先通过BeautifulSoup定位到包含
window.__PRELOADED_STATE__的script标签,缩小匹配范围,避免HTML中其他内容干扰 - 使用精准正则匹配
JSON.parse()括号内的JSON字符串,处理转义字符后用json.loads转为字典
示例代码
from bs4 import BeautifulSoup import re import json # 替换为你的HTML内容 html = """ <html> <body> <script> window.__PRELOADED_STATE__ = JSON.parse('{"userInfo": {"name": "Behrouz", "id": 123}, "settings": {"theme": "dark"}}'); </script> </body> </html> """ # 定位目标script标签 soup = BeautifulSoup(html, "html.parser") target_script = None for script in soup.find_all("script"): if script.string and "window.__PRELOADED_STATE__" in script.string: target_script = script.string break if target_script: # 匹配JSON.parse内的内容,支持多行字符串 match = re.search( r"window\.__PRELOADED_STATE__\s*=\s*JSON\.parse\((.*?)\);", target_script, re.DOTALL ) if match: raw_json_str = match.group(1).strip() # 去除外层包裹的单/双引号 if (raw_json_str.startswith("'") and raw_json_str.endswith("'")) or \ (raw_json_str.startswith('"') and raw_json_str.endswith('"')): raw_json_str = raw_json_str[1:-1] # 还原转义的引号,避免JSON解析失败 processed_json_str = raw_json_str.replace(r"\'", "'").replace(r'\"', '"') # 解析为字典 try: preloaded_state = json.loads(processed_json_str) print("解析成功,结果类型:", type(preloaded_state)) print("预加载数据:", preloaded_state) except json.JSONDecodeError as e: print(f"JSON解析失败: {str(e)}") else: print("未匹配到JSON.parse中的内容") else: print("未找到包含预加载状态的script标签")
关键注意点
- 必须使用
re.DOTALL正则标志,否则无法匹配多行的JSON字符串 - 要处理JSON字符串外层的引号包裹,以及字符串内部的转义引号,否则
json.loads会报错 - 先通过BeautifulSoup筛选script标签,避免直接在整个HTML中匹配正则时出现误匹配
内容的提问来源于stack exchange,提问作者Behrouz Hedayati
相关产品推荐
相关产品推荐

