如何用Python从HTML中提取JSON对象的lineItemId?
提取window.__EP对象中lineItemId的解决方案
核心问题分析
BeautifulSoup负责解析HTML结构,但window.__EP是嵌入在script标签里的JavaScript对象,无法直接用HTML解析方法提取,得先定位到目标script标签,再提取并转换JS对象为可解析的JSON格式。
分步解决代码
- 定位目标script标签
遍历页面所有script标签,筛选出包含window.__EP的标签内容:
from bs4 import BeautifulSoup import re import json # 假设html是你获取到的页面源码 soup = BeautifulSoup(html, 'html.parser') target_script = None for script in soup.find_all('script'): if script.string and 'window.__EP' in script.string: target_script = script.string break
- 提取并转换JS对象为JSON
用正则匹配出window.__EP对应的对象内容,再修复JS语法使其符合JSON规范:
if target_script: # 匹配window.__EP = { ... }的结构,支持换行 match = re.search(r'window.__EP\s*=\s*({.*?});', target_script, re.DOTALL) if match: js_obj_str = match.group(1) # 修复JS到JSON的语法差异 json_str = js_obj_str.replace("'", "\"") # 单引号转双引号 json_str = json_str.replace("undefined", "null") # undefined替换为null # 若有其他JS语法(如函数),需额外处理,这里假设无复杂结构 try: ep_data = json.loads(json_str) # 提取目标lineItemId product_line_items = ep_data.get('storeInitialState', {}).get('cart', {}).get('productLineItems', []) line_item_ids = [item.get('lineItemId') for item in product_line_items if item.get('lineItemId')] print(line_item_ids) except json.JSONDecodeError as e: print(f"JSON解析失败: {e}") else: print("未匹配到window.__EP对象") else: print("未找到包含window.__EP的script标签")
常见坑点处理
- 如果页面script内容是压缩后的(无换行),正则的
re.DOTALL依然有效,但要确保匹配的{.*?}能覆盖整个对象(若对象末尾有其他代码,可能需要调整正则的结束匹配条件,比如匹配到};或其他标识) - 若JS对象里有
NaN、函数定义等,需要额外替换或剔除,比如把NaN换成null,用正则去掉function\s*\(.*?\)\s*{.*?}这类函数内容
内容的提问来源于stack exchange,提问作者James Carlton
相关产品推荐
相关产品推荐

