You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取网页script标签内的数据?

提取Daraz站点商品名称的实现方法

Daraz站点的商品列表数据未直接渲染在HTML可见标签中,而是预存在<script>标签内的window.__INITIAL_STATE__JS全局变量中,数据为标准JSON格式,可按以下方案实现提取:

完整可运行代码

import requests
from bs4 import BeautifulSoup
import json

# 加请求头避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

searchf = input("Enter the product you want: ")

url =f"https://www.daraz.com.np/catalog/?q={searchf}&_keyori=ss&from=input&spm=a2a0e.searchlist.search.go.12ec45adZnn1MP"

r = requests.get(url, headers=headers)
contents = r.content

soup = BeautifulSoup(contents,'html.parser')

# 定位包含商品数据的script标签
target_script = None
for script in soup.find_all("script"):
    script_content = script.string
    if script_content and "window.__INITIAL_STATE__" in script_content:
        target_script = script_content
        break

# 提取解析JSON数据
if target_script:
    # 截取JSON字符串,移除前后的JS代码
    json_raw = target_script.split("window.__INITIAL_STATE__ = ")[1].rsplit(";", 1)[0]
    page_data = json.loads(json_raw)
    # 提取商品列表
    product_list = page_data["mods"]["listItems"]
    # 输出所有商品名称
    for product in product_list:
        print(product["name"])

核心逻辑说明

  • 新增User-Agent请求头,绕过站点基础反爬拦截,避免返回无效空内容
  • 遍历所有<script>标签,匹配到包含window.__INITIAL_STATE__变量的目标标签
  • 截取变量对应的JSON字符串,移除多余的JS语法字符后用json库解析为Python字典
  • 从解析后的字典结构中定位listItems商品数组,遍历取出所有商品的name字段即可得到所需的商品名称

内容的提问来源于stack exchange,提问作者A-d-ash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 03:06:00