You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup报错'NoneType' object has no attribute 'find_all'求助

爬取Wuzzuf.net招聘信息时遭遇AttributeError报错

作为网页抓取新手,爬取wuzzuf.net招聘信息时执行代码出现以下报错:

AttributeError: 'NoneType' object has no attribute 'find_all'

报错发生在获取sal_span的链式调用语句中,此前用类似链式调用方法获取years变量时运行正常。

我的代码

from bs4 import BeautifulSoup
import requests
import csv
from itertools import zip_longest
import re

jobTitle = []
companyName = []
location = []
employmentReq = []
exp = []
startTime = []
links = []
salary = []
experience = []
pageNum = 0
result = requests.get(f"https://wuzzuf.net/search/jobs/?a=hpb&q=web&start={pageNum}")

src = result.content

soup = BeautifulSoup(src, "html.parser")

jobTitles = soup.find_all("h2", {"class": "css-m604qf"})
companyNames = soup.find_all("a", {"class": "css-17s97q8"})
locations = soup.find_all("span", {"class": "css-5wys0k"})
employment = soup.find_all("div", {"class": "css-1lh32fc"})

for i in range(len(jobTitles)):
    jobTitle.append(jobTitles[i].text)
    links.append("https://wuzzuf.net" + jobTitles[i].find("a").attrs['href'])
    companyName.append(companyNames[i].text)
    location.append(locations[i].text)
    employmentReq.append(employment[i].text)
    years = re.sub(r'[^0-9-]', '', soup.find_all("div", {"class": "css-y4udm8"})[i].find_all("div")[1].find_all("span")[0].text)
    experience.append(years)

for link in links:
    result = requests.get(link)
    src = result.content
    soup = BeautifulSoup(src, "html.parser")
    a = soup.find("main")
    print(a)
    b=a.find("section",{"class":"css-3kx5e2"})
    print(b)
    c =b.find_all("div")
    print(c)
    d =c.find_all("span")
    print(d)

    #if sal_span != "Confidential":
    #    salaries = re.sub(r'E.*$', '', sal_span)
    #else:
    #    salaries = sal_span
    #salary.append(salaries)

#print(salary)
#fileList = [jobTitle, companyName, location, employmentReq, links, experience, salary]
#exported = zip_longest(*fileList)
#with open("D:\T1t4nProject\python\wuzzuf.csv", "w") as excel_sheet:
#    wr = csv.writer(excel_sheet)
#    wr.writerow(["job title", "company name", "location", "full or part time", "links", "Years of experience", "Salary"])
#    wr.writerows(exported)

报错详情

sal_span =soup.find("main").find("section",{"class":"css-3kx5e2"}).find_all("div")[3].find_all("span")[1].find("span").text
AttributeError: 'NoneType' object has no attribute 'find_all'

注:years变量是从搜索结果页面获取的,而薪资信息需要进入每个职位详情页抓取。


问题原因

链式调用中某一个节点返回了None(比如找不到<main>标签、目标section,或者按索引取的div/span不存在),导致后续调用find_all时触发报错。这种情况通常是因为:

  • 部分职位详情页没有公开薪资信息,页面结构和预期不一致
  • 网站页面结构更新,原来的类名或节点位置发生变化
  • 依赖固定索引(如find_all("div")[3])定位元素,一旦页面元素顺序调整就会失效

解决方案

1. 分步检查节点,避免链式调用

不要一次性把所有find/find_all连在一起,逐个检查每个节点是否存在,不存在时给薪资赋值为"N/A"或其他默认值。

2. 使用更灵活的选择器定位薪资

避免依赖固定索引,改用文本匹配或更精准的类选择器定位薪资相关元素。

修改后的代码片段(薪资抓取部分)

for link in links:
    result = requests.get(link)
    src = result.content
    soup = BeautifulSoup(src, "html.parser")
    
    # 分步获取节点,逐个检查是否存在
    main = soup.find("main")
    if not main:
        salary.append("N/A")
        continue
    
    section = main.find("section", {"class": "css-3kx5e2"})
    if not section:
        salary.append("N/A")
        continue
    
    # 通过文本匹配找到包含"Salary"的标签,再定位对应的薪资内容
    salary_label = section.find("div", string=re.compile(r"Salary"))
    if salary_label:
        salary_content = salary_label.find_next_sibling("div")
        if salary_content:
            sal_span = salary_content.find("span")
            if sal_span:
                sal_text = sal_span.text.strip()
                # 处理薪资格式
                if sal_text != "Confidential":
                    salaries = re.sub(r'E.*$', '', sal_text)
                else:
                    salaries = sal_text
                salary.append(salaries)
            else:
                salary.append("N/A")
        else:
            salary.append("N/A")
    else:
        salary.append("N/A")

额外建议

  • 加入请求头(如User-Agent)模拟浏览器请求,避免被网站反爬机制拦截
  • 给请求添加延时(如time.sleep(1)),减轻目标网站服务器压力
  • 对所有可能返回None的节点都做存在性检查,提升代码稳定性

内容的提问来源于stack exchange,提问作者Titan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 19:25:28