Python网页抓取:如何从JSON-LD的department数组提取name与url
提取JSON-LD中的部门名称与链接
先修正你原代码里的问题,再一步步实现提取需求:
1. 修正基础代码
原代码存在变量名错误、选择器写法错误以及无效操作,调整为可运行的基础代码:
import requests from bs4 import BeautifulSoup import json # 发起请求并获取响应 res = requests.get("http://www.ucdenver.edu/pages/ucdwelcomepage.aspx") # 解析HTML soup = BeautifulSoup(res.content, 'html5lib') # 筛选指定类型的script标签 scripts = soup.select('script[type="application/ld+json"]')
2. 解析JSON-LD并提取目标字段
每个符合条件的script标签内的文本是JSON格式数据,转成Python字典后即可遍历提取字段:
for script in scripts: # 提取script标签内的文本内容 json_text = script.string if not json_text: continue # 跳过空内容的标签 # 解析JSON为Python字典 try: data = json.loads(json_text) except json.JSONDecodeError: continue # 跳过解析失败的内容 # 提取department数组中的字段 if "department" in data: departments = data["department"] # 遍历每个部门对象 for dept in departments: # 用get方法取值,避免字段缺失报错 dept_name = dept.get("name", "无名称") dept_url = dept.get("url", "无链接") print(f"部门名称: {dept_name}") print(f"部门链接: {dept_url}") print("---")
关键说明
script.string:获取script标签包裹的文本,这是JSON数据的来源json.loads():将JSON字符串转换为Python字典/列表,方便后续取值dept.get(key, 默认值):容错处理,避免因字段缺失导致程序报错
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

