You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页抓取:如何从JSON-LD的department数组提取name与url

提取JSON-LD中的部门名称与链接

先修正你原代码里的问题,再一步步实现提取需求:

1. 修正基础代码

原代码存在变量名错误、选择器写法错误以及无效操作,调整为可运行的基础代码:

import requests
from bs4 import BeautifulSoup
import json

# 发起请求并获取响应
res = requests.get("http://www.ucdenver.edu/pages/ucdwelcomepage.aspx")
# 解析HTML
soup = BeautifulSoup(res.content, 'html5lib')
# 筛选指定类型的script标签
scripts = soup.select('script[type="application/ld+json"]')

2. 解析JSON-LD并提取目标字段

每个符合条件的script标签内的文本是JSON格式数据,转成Python字典后即可遍历提取字段:

for script in scripts:
    # 提取script标签内的文本内容
    json_text = script.string
    if not json_text:
        continue  # 跳过空内容的标签
    # 解析JSON为Python字典
    try:
        data = json.loads(json_text)
    except json.JSONDecodeError:
        continue  # 跳过解析失败的内容
    
    # 提取department数组中的字段
    if "department" in data:
        departments = data["department"]
        # 遍历每个部门对象
        for dept in departments:
            # 用get方法取值,避免字段缺失报错
            dept_name = dept.get("name", "无名称")
            dept_url = dept.get("url", "无链接")
            print(f"部门名称: {dept_name}")
            print(f"部门链接: {dept_url}")
            print("---")

关键说明

  • script.string:获取script标签包裹的文本,这是JSON数据的来源
  • json.loads():将JSON字符串转换为Python字典/列表,方便后续取值
  • dept.get(key, 默认值):容错处理,避免因字段缺失导致程序报错

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 08:50:26