You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取网站JSON触发KeyError错误,求故障排查思路

问题原因
  • 页面中存在多个application/ld+json类型的脚本标签,并非所有标签对应的JSON都包含itemListElement字段,你的代码直接通过jsn['itemListElement']取值,遇到没有该字段的JSON时就会触发KeyError。
  • 你错误默认所有匹配到的脚本都符合包含职位列表的结构,但实际上页面可能还有其他用途的LD+JSON数据(比如网站自身的结构化信息)。
解决建议
  1. 增加字段存在性检查:在访问itemListElement前,先判断该字段是否存在于当前JSON对象中,跳过不符合结构的脚本。
  2. 修正DataFrame赋值逻辑:原代码中df['URL'] = ...会覆盖整列数据,应该用逐行添加的方式填充职位信息。

修正后的代码示例

import requests
from bs4 import BeautifulSoup
import json
import pandas as pd

headers = {
    "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:90.0) Gecko/20100101 Firefox/90.0"
}
url = "https://employment.ucsd.edu/jobs?page_size=250&page_number=1&keyword=clinical%20lab%20scientist&location_city" \
      "=Remote&location_city=San%20Diego&location_city=Encinitas&location_city=Murrieta&location_city=La%20Jolla" \
      "&location_city=Not%20Specified&location_city=Vista&sort_by=score&sort_order=DESC "
request = requests.get(url, headers=headers)
response = BeautifulSoup(request.text, "html.parser")
all_data = response.find_all("script", {"type": "application/ld+json"})
df = pd.DataFrame(columns=("Title", "Department", "Salary Range", "Appointment Percent", "URL"))

for data in all_data:
    try:
        jsn = json.loads(data.string)
        # 先检查itemListElement是否存在,不存在则跳过当前脚本
        if 'itemListElement' not in jsn:
            continue
        # 遍历职位列表
        for item in jsn['itemListElement']:
            # 根据实际JSON结构提取字段,这里是示例逻辑
            job_data = {
                "Title": item.get('name', 'N/A'),
                "URL": item.get('url', 'N/A'),
                "Department": item.get('hiringOrganization', {}).get('name', 'N/A'),
                "Salary Range": "",  # 需根据实际返回的JSON结构补充提取逻辑
                "Appointment Percent": ""
            }
            # 逐行添加到DataFrame
            df = df._append(job_data, ignore_index=True)
    except json.JSONDecodeError:
        # 跳过解析失败的脚本
        continue

print(df)

额外提示

  • 可以取消注释原代码里的print(json.dumps(jsn, indent=4)),打印所有匹配到的JSON结构,确认哪些脚本包含职位数据,以及字段的具体路径。
  • 使用dict.get()方法可以避免KeyError,还能设置默认值(比如item.get('name', 'N/A'))。

内容的提问来源于stack exchange,提问作者Baby Yoda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 13:45:31