You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup精准查找类?爬取招聘网站遇NoneType错误求助

解决BeautifulSoup爬取联合航空招聘网站的AttributeError问题

问题根源

你遇到的AttributeError是因为soup.find(class_="jobTitle")返回了None——目标网站的职位标题元素根本没有使用jobTitle这个类名,或者请求没有正确获取到页面内容。

分步解决方法

  • 先确认请求有效性
    很多网站会拦截无浏览器标识的请求,先给请求加User-Agent头,同时检查响应状态码:

    import requests
    from bs4 import BeautifulSoup
    
    # 模拟浏览器请求头
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    page = requests.get("https://careers.united.com/job-search-results/", headers=headers)
    print(page.status_code)  # 输出200才代表请求成功
    
  • 定位正确的元素类名
    打开目标网站,右键点击任意职位标题选择「检查」,查看元素的实际类名(比如可能是job-title、card-title这类命名)。找到后替换代码中的类名,比如实际类名是job-title的话:

    soup = BeautifulSoup(page.content, "html.parser")
    # 用find_all获取所有职位标题,而不是单个
    job_titles = soup.find_all(class_="job-title")
    for title in job_titles:
        print(title.get_text(strip=True))
    
  • 处理动态渲染内容
    如果检查静态HTML里找不到职位元素,说明内容是JavaScript动态加载的。这时需要用Selenium获取渲染后的页面:

    from selenium import webdriver
    from bs4 import BeautifulSoup
    
    driver = webdriver.Chrome()
    driver.get("https://careers.united.com/job-search-results/")
    # 等待页面加载完成(可根据实际情况加显式等待)
    soup = BeautifulSoup(driver.page_source, "html.parser")
    job_titles = soup.find_all(class_="实际找到的类名")
    for title in job_titles:
        print(title.get_text(strip=True))
    driver.quit()
    
  • 避免NoneType报错的安全写法
    不管怎样,先判断元素是否存在再调用方法:

    result = soup.find(class_="正确类名")
    if result:
        print(result.prettify())
    else:
        print("未找到目标元素,请检查类名或页面结构")
    

内容的提问来源于stack exchange,提问作者MalTechh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 00:35:33