You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取含window.__additionalDataLoaded的脚本中的JSON

提取Instagram嵌入页中window.__additionalDataLoaded的JSON数据

核心思路

目标脚本的核心特征是包含window.__additionalDataLoaded('extra'字符串,先精准定位该脚本标签,再提取其中的JSON结构内容。

具体实现步骤

  • 步骤1:定位目标脚本标签
    放弃仅通过type="text/javascript"筛选的方式,直接匹配脚本内容中包含目标特征字符串的标签。以Python为例,用BeautifulSoup实现:

    from bs4 import BeautifulSoup
    import requests
    import re
    import json
    
    url = "https://www.instagram.com/p/CgKjwd-JR3p/embed/captioned/"
    response = requests.get(url)
    soup = BeautifulSoup(response.text, "html.parser")
    
    target_script = None
    # 遍历所有script标签,找到包含目标特征的内容
    for script in soup.find_all("script"):
        if script.string and "window.__additionalDataLoaded('extra'" in script.string:
            target_script = script.string
            break
    
  • 步骤2:提取并解析纯净JSON
    目标脚本的结构为window.__additionalDataLoaded('extra', { ...JSON内容... });,用正则匹配捕获从第一个{到最后一个}的完整内容,同时开启多行匹配模式处理换行:

    if target_script:
        # 匹配完整的JSON部分,re.DOTALL允许.匹配换行符
        json_match = re.search(r'window.__additionalDataLoaded\(\'extra\',\s*({.*})\s*\);', target_script, re.DOTALL)
        if json_match:
            json_content = json_match.group(1)
            # 解析JSON为字典
            parsed_data = json.loads(json_content)
            # 输出或使用解析后的数据
            print(parsed_data)
    

注意事项

  • 若页面结构有微小调整,只需微调正则中匹配函数调用的部分,核心逻辑保持不变;
  • json.loads会自动处理JSON中的转义字符,无需手动干预。

内容的提问来源于stack exchange,提问作者Hossam Al-din Hassan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 15:27:20