爬取网站JSON触发KeyError错误,求故障排查思路
问题原因
- 页面中存在多个
application/ld+json类型的脚本标签,并非所有标签对应的JSON都包含itemListElement字段,你的代码直接通过jsn['itemListElement']取值,遇到没有该字段的JSON时就会触发KeyError。 - 你错误默认所有匹配到的脚本都符合包含职位列表的结构,但实际上页面可能还有其他用途的LD+JSON数据(比如网站自身的结构化信息)。
解决建议
- 增加字段存在性检查:在访问
itemListElement前,先判断该字段是否存在于当前JSON对象中,跳过不符合结构的脚本。 - 修正DataFrame赋值逻辑:原代码中
df['URL'] = ...会覆盖整列数据,应该用逐行添加的方式填充职位信息。
修正后的代码示例
import requests from bs4 import BeautifulSoup import json import pandas as pd headers = { "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:90.0) Gecko/20100101 Firefox/90.0" } url = "https://employment.ucsd.edu/jobs?page_size=250&page_number=1&keyword=clinical%20lab%20scientist&location_city" \ "=Remote&location_city=San%20Diego&location_city=Encinitas&location_city=Murrieta&location_city=La%20Jolla" \ "&location_city=Not%20Specified&location_city=Vista&sort_by=score&sort_order=DESC " request = requests.get(url, headers=headers) response = BeautifulSoup(request.text, "html.parser") all_data = response.find_all("script", {"type": "application/ld+json"}) df = pd.DataFrame(columns=("Title", "Department", "Salary Range", "Appointment Percent", "URL")) for data in all_data: try: jsn = json.loads(data.string) # 先检查itemListElement是否存在,不存在则跳过当前脚本 if 'itemListElement' not in jsn: continue # 遍历职位列表 for item in jsn['itemListElement']: # 根据实际JSON结构提取字段,这里是示例逻辑 job_data = { "Title": item.get('name', 'N/A'), "URL": item.get('url', 'N/A'), "Department": item.get('hiringOrganization', {}).get('name', 'N/A'), "Salary Range": "", # 需根据实际返回的JSON结构补充提取逻辑 "Appointment Percent": "" } # 逐行添加到DataFrame df = df._append(job_data, ignore_index=True) except json.JSONDecodeError: # 跳过解析失败的脚本 continue print(df)
额外提示
- 可以取消注释原代码里的
print(json.dumps(jsn, indent=4)),打印所有匹配到的JSON结构,确认哪些脚本包含职位数据,以及字段的具体路径。 - 使用
dict.get()方法可以避免KeyError,还能设置默认值(比如item.get('name', 'N/A'))。
内容的提问来源于stack exchange,提问作者Baby Yoda
相关产品推荐
相关产品推荐

