Python网页抓取求助:提取<script>中var的JSON数据并生成数据集
网页抓取问题:提取bilancio_tree中的收入数据
需求说明
需要从目标网站抓取<script>标签内var bilancio_tree中的JSON数据,提取各项收入的abs和pc值,生成结构化数据集。目标数据示例如下:
var bilancio_tree = [{"slug": "pcox-quadro-2-11", "label": "Totale generale delle Entrate", "values": [{"2021": {"abs": 1659238.91, "pc": 3463.96432150313}}, {"2022": {"abs": 0.0, "pc": 0.0}}, {"2023": {"abs": 0.0, "pc": 0.0}}], "children": []}, ...];
期望生成的结构化数据集格式:
| Totale generale delle Entrate Total | Totale generale delle Entrate PC |
|---|---|
| 1659238.91 | 3463.96 |
原脚本问题分析
原脚本存在以下问题导致无法正常运行:
- 硬编码使用
soup.find_all("script")[19]获取脚本标签,页面结构变化时会直接失效 - 使用
re.match()仅从字符串开头匹配,而var bilancio_tree大概率不在脚本内容的起始位置 - 未处理匹配失败的情况,若找不到目标数据会直接抛出异常中断程序
修复后的Python脚本
import requests from bs4 import BeautifulSoup import json import re URL = "https://openbilanci.it/armonizzati/bilanci/veglio-comune-bi/entrate/dettaglio?year=2021&type=preventivo" r = requests.get(URL) soup = BeautifulSoup(r.content, 'html.parser') # 遍历所有script标签,定位包含目标数据的脚本 target_script = None for script in soup.find_all("script"): if script.string and 'var bilancio_tree' in script.string: target_script = script.string break if not target_script: print("未找到目标数据") else: # 匹配整个文本中的目标JSON片段 match = re.search(r'var bilancio_tree = (.*?);', target_script, re.DOTALL) if match: try: bilancio_data = json.loads(match.group(1)) # 格式化输出为Markdown表格 print("| 收入项目 | 绝对值(abs) | 占比(pc) |") print("|----------|-------------|----------|") for item in bilancio_data: # 获取2021年的数据(可按需修改年份) year_data = next((val for val in item['values'] if '2021' in val), None) if year_data: abs_val = year_data['2021']['abs'] pc_val = round(year_data['2021']['pc'], 2) print(f"| {item['label']} | {abs_val} | {pc_val} |") except json.JSONDecodeError as e: print(f"JSON解析错误: {e}") else: print("未匹配到目标JSON数据")
脚本说明
- 动态定位脚本:遍历所有
<script>标签,通过关键字匹配找到目标脚本,避免固定索引依赖 - 正则匹配优化:使用
re.search()结合re.DOTALL标志,支持匹配多行文本中的JSON片段 - 异常处理:增加数据未找到、JSON解析失败的异常捕获,提升脚本稳定性
- 结构化输出:提取指定年份的
abs和pc值,自动格式化为易读的Markdown表格
内容的提问来源于stack exchange,提问作者Pepa
相关产品推荐
相关产品推荐

