You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup和正则提取JavaScript中json变量的值

提取JavaScript中json变量值的实现方案

没问题,我来帮你搞定这个需求!要提取这段JS代码里的json变量值,我们可以结合BeautifulSoup定位目标script标签,再用正则表达式精准匹配变量内容,下面是具体的实现步骤和代码:

步骤说明

  1. 解析HTML获取script内容:先用BeautifulSoup从HTML(或代码片段)里筛选出包含目标变量的script标签;
  2. 正则匹配变量值:用正则表达式捕获json变量赋值语句里的JSON字符串内容。

完整代码示例

from bs4 import BeautifulSoup
import re

# 模拟包含目标script标签的HTML内容(如果是从网页爬取,这里可以替换为requests.get获取的响应文本)
html_content = """
<script type="text/javascript"> $(document).ready(function() { var dataJson = ""; var grid = $("#projectTable"); var pagerDiv = $("#pagingDiv").attr('id'); var gridId = $("#projectTable").attr('id'); var fileType = '1'; var quarterId = '78'; var totalPages = 0; var json = '{"total":3,"records":44,"page":1}' </script>
"""

# 1. 用BeautifulSoup解析HTML,找到目标script标签
soup = BeautifulSoup(html_content, 'html.parser')
# 筛选type为text/javascript且内容包含"var json ="的script标签
target_script = soup.find('script', {'type': 'text/javascript'}, string=re.compile(r'var json ='))

if target_script:
    # 2. 提取script标签内的文本内容
    script_text = target_script.string
    # 3. 用正则匹配json变量的值(匹配单引号包裹的JSON字符串)
    match = re.search(r"var json = '([\s\S]*?)'", script_text)
    if match:
        json_value = match.group(1)
        print("提取到的json变量值:")
        print(json_value)
        # 如果需要转为Python字典,可以用json.loads
        import json
        json_dict = json.loads(json_value)
        print("\n转为Python字典后:")
        print(json_dict)
    else:
        print("未找到json变量的赋值语句")
else:
    print("未找到包含目标变量的script标签")

代码解释

  • BeautifulSoup筛选:通过find方法结合属性和正则匹配,精准定位包含var json =的script标签,避免匹配其他无关的script;
  • 正则表达式:r"var json = '([\s\S]*?)'" 的作用是匹配以var json = '开头、以'结尾的内容,[\s\S]*?用来匹配任意字符(包括换行)且采用非贪婪模式,确保只捕获当前变量的赋值内容;
  • 可选扩展:如果需要把提取到的JSON字符串转为Python字典,可以用json.loads()方法直接解析。

内容的提问来源于stack exchange,提问作者sudmat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 14:52:37