BeautifulSoup报错'NoneType' object has no attribute 'find_all'求助
爬取Wuzzuf.net招聘信息时遭遇AttributeError报错
作为网页抓取新手,爬取wuzzuf.net招聘信息时执行代码出现以下报错:
AttributeError: 'NoneType' object has no attribute 'find_all'
报错发生在获取sal_span的链式调用语句中,此前用类似链式调用方法获取years变量时运行正常。
我的代码
from bs4 import BeautifulSoup import requests import csv from itertools import zip_longest import re jobTitle = [] companyName = [] location = [] employmentReq = [] exp = [] startTime = [] links = [] salary = [] experience = [] pageNum = 0 result = requests.get(f"https://wuzzuf.net/search/jobs/?a=hpb&q=web&start={pageNum}") src = result.content soup = BeautifulSoup(src, "html.parser") jobTitles = soup.find_all("h2", {"class": "css-m604qf"}) companyNames = soup.find_all("a", {"class": "css-17s97q8"}) locations = soup.find_all("span", {"class": "css-5wys0k"}) employment = soup.find_all("div", {"class": "css-1lh32fc"}) for i in range(len(jobTitles)): jobTitle.append(jobTitles[i].text) links.append("https://wuzzuf.net" + jobTitles[i].find("a").attrs['href']) companyName.append(companyNames[i].text) location.append(locations[i].text) employmentReq.append(employment[i].text) years = re.sub(r'[^0-9-]', '', soup.find_all("div", {"class": "css-y4udm8"})[i].find_all("div")[1].find_all("span")[0].text) experience.append(years) for link in links: result = requests.get(link) src = result.content soup = BeautifulSoup(src, "html.parser") a = soup.find("main") print(a) b=a.find("section",{"class":"css-3kx5e2"}) print(b) c =b.find_all("div") print(c) d =c.find_all("span") print(d) #if sal_span != "Confidential": # salaries = re.sub(r'E.*$', '', sal_span) #else: # salaries = sal_span #salary.append(salaries) #print(salary) #fileList = [jobTitle, companyName, location, employmentReq, links, experience, salary] #exported = zip_longest(*fileList) #with open("D:\T1t4nProject\python\wuzzuf.csv", "w") as excel_sheet: # wr = csv.writer(excel_sheet) # wr.writerow(["job title", "company name", "location", "full or part time", "links", "Years of experience", "Salary"]) # wr.writerows(exported)
报错详情
sal_span =soup.find("main").find("section",{"class":"css-3kx5e2"}).find_all("div")[3].find_all("span")[1].find("span").text AttributeError: 'NoneType' object has no attribute 'find_all'
注:years变量是从搜索结果页面获取的,而薪资信息需要进入每个职位详情页抓取。
问题原因
链式调用中某一个节点返回了None(比如找不到<main>标签、目标section,或者按索引取的div/span不存在),导致后续调用find_all时触发报错。这种情况通常是因为:
- 部分职位详情页没有公开薪资信息,页面结构和预期不一致
- 网站页面结构更新,原来的类名或节点位置发生变化
- 依赖固定索引(如
find_all("div")[3])定位元素,一旦页面元素顺序调整就会失效
解决方案
1. 分步检查节点,避免链式调用
不要一次性把所有find/find_all连在一起,逐个检查每个节点是否存在,不存在时给薪资赋值为"N/A"或其他默认值。
2. 使用更灵活的选择器定位薪资
避免依赖固定索引,改用文本匹配或更精准的类选择器定位薪资相关元素。
修改后的代码片段(薪资抓取部分)
for link in links: result = requests.get(link) src = result.content soup = BeautifulSoup(src, "html.parser") # 分步获取节点,逐个检查是否存在 main = soup.find("main") if not main: salary.append("N/A") continue section = main.find("section", {"class": "css-3kx5e2"}) if not section: salary.append("N/A") continue # 通过文本匹配找到包含"Salary"的标签,再定位对应的薪资内容 salary_label = section.find("div", string=re.compile(r"Salary")) if salary_label: salary_content = salary_label.find_next_sibling("div") if salary_content: sal_span = salary_content.find("span") if sal_span: sal_text = sal_span.text.strip() # 处理薪资格式 if sal_text != "Confidential": salaries = re.sub(r'E.*$', '', sal_text) else: salaries = sal_text salary.append(salaries) else: salary.append("N/A") else: salary.append("N/A") else: salary.append("N/A")
额外建议
- 加入请求头(如
User-Agent)模拟浏览器请求,避免被网站反爬机制拦截 - 给请求添加延时(如
time.sleep(1)),减轻目标网站服务器压力 - 对所有可能返回
None的节点都做存在性检查,提升代码稳定性
内容的提问来源于stack exchange,提问作者Titan
相关产品推荐
相关产品推荐

