You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从HTML中提取JSON对象的lineItemId?

提取window.__EP对象中lineItemId的解决方案

核心问题分析

BeautifulSoup负责解析HTML结构,但window.__EP是嵌入在script标签里的JavaScript对象,无法直接用HTML解析方法提取,得先定位到目标script标签,再提取并转换JS对象为可解析的JSON格式。

分步解决代码

  1. 定位目标script标签
    遍历页面所有script标签,筛选出包含window.__EP的标签内容:
from bs4 import BeautifulSoup
import re
import json

# 假设html是你获取到的页面源码
soup = BeautifulSoup(html, 'html.parser')
target_script = None
for script in soup.find_all('script'):
    if script.string and 'window.__EP' in script.string:
        target_script = script.string
        break
  1. 提取并转换JS对象为JSON
    用正则匹配出window.__EP对应的对象内容,再修复JS语法使其符合JSON规范:
if target_script:
    # 匹配window.__EP = { ... }的结构,支持换行
    match = re.search(r'window.__EP\s*=\s*({.*?});', target_script, re.DOTALL)
    if match:
        js_obj_str = match.group(1)
        # 修复JS到JSON的语法差异
        json_str = js_obj_str.replace("'", "\"")  # 单引号转双引号
        json_str = json_str.replace("undefined", "null")  # undefined替换为null
        # 若有其他JS语法(如函数),需额外处理,这里假设无复杂结构
        try:
            ep_data = json.loads(json_str)
            # 提取目标lineItemId
            product_line_items = ep_data.get('storeInitialState', {}).get('cart', {}).get('productLineItems', [])
            line_item_ids = [item.get('lineItemId') for item in product_line_items if item.get('lineItemId')]
            print(line_item_ids)
        except json.JSONDecodeError as e:
            print(f"JSON解析失败: {e}")
    else:
        print("未匹配到window.__EP对象")
else:
    print("未找到包含window.__EP的script标签")

常见坑点处理

  • 如果页面script内容是压缩后的(无换行),正则的re.DOTALL依然有效,但要确保匹配的{.*?}能覆盖整个对象(若对象末尾有其他代码,可能需要调整正则的结束匹配条件,比如匹配到};或其他标识)
  • 若JS对象里有NaN、函数定义等,需要额外替换或剔除,比如把NaN换成null,用正则去掉function\s*\(.*?\)\s*{.*?}这类函数内容

内容的提问来源于stack exchange,提问作者James Carlton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 09:36:28