You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页抓取求助:提取<script>中var的JSON数据并生成数据集

网页抓取问题:提取bilancio_tree中的收入数据

需求说明

需要从目标网站抓取<script>标签内var bilancio_tree中的JSON数据,提取各项收入的abs和pc值,生成结构化数据集。目标数据示例如下:

var bilancio_tree = [{"slug": "pcox-quadro-2-11", "label": "Totale generale delle Entrate", "values": [{"2021": {"abs": 1659238.91, "pc": 3463.96432150313}}, {"2022": {"abs": 0.0, "pc": 0.0}}, {"2023": {"abs": 0.0, "pc": 0.0}}], "children": []}, ...];

期望生成的结构化数据集格式:

Totale generale delle Entrate TotalTotale generale delle Entrate PC
1659238.913463.96

原脚本问题分析

原脚本存在以下问题导致无法正常运行:

  • 硬编码使用soup.find_all("script")[19]获取脚本标签,页面结构变化时会直接失效
  • 使用re.match()仅从字符串开头匹配,而var bilancio_tree大概率不在脚本内容的起始位置
  • 未处理匹配失败的情况,若找不到目标数据会直接抛出异常中断程序

修复后的Python脚本

import requests
from bs4 import BeautifulSoup
import json
import re

URL = "https://openbilanci.it/armonizzati/bilanci/veglio-comune-bi/entrate/dettaglio?year=2021&type=preventivo"
r = requests.get(URL)
soup = BeautifulSoup(r.content, 'html.parser')

# 遍历所有script标签,定位包含目标数据的脚本
target_script = None
for script in soup.find_all("script"):
    if script.string and 'var bilancio_tree' in script.string:
        target_script = script.string
        break

if not target_script:
    print("未找到目标数据")
else:
    # 匹配整个文本中的目标JSON片段
    match = re.search(r'var bilancio_tree = (.*?);', target_script, re.DOTALL)
    if match:
        try:
            bilancio_data = json.loads(match.group(1))
            # 格式化输出为Markdown表格
            print("| 收入项目 | 绝对值(abs) | 占比(pc) |")
            print("|----------|-------------|----------|")
            for item in bilancio_data:
                # 获取2021年的数据(可按需修改年份)
                year_data = next((val for val in item['values'] if '2021' in val), None)
                if year_data:
                    abs_val = year_data['2021']['abs']
                    pc_val = round(year_data['2021']['pc'], 2)
                    print(f"| {item['label']} | {abs_val} | {pc_val} |")
        except json.JSONDecodeError as e:
            print(f"JSON解析错误: {e}")
    else:
        print("未匹配到目标JSON数据")

脚本说明

  1. 动态定位脚本:遍历所有<script>标签,通过关键字匹配找到目标脚本,避免固定索引依赖
  2. 正则匹配优化:使用re.search()结合re.DOTALL标志,支持匹配多行文本中的JSON片段
  3. 异常处理:增加数据未找到、JSON解析失败的异常捕获,提升脚本稳定性
  4. 结构化输出:提取指定年份的abs和pc值,自动格式化为易读的Markdown表格

内容的提问来源于stack exchange,提问作者Pepa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:00:45