Python BeautifulSoup爬虫find方法触发NoneType错误求解
问题定位
- 未校验外层容器存在性:代码直接对
find("div",{"class":"css-1t5f0fr"})的返回结果取.ul/.p属性,若该容器不存在会直接得到None,后续调用方法就会触发NoneType报错。 - 相对路径未补全:抓取到的职位详情链接是相对路径,直接调用
requests.get()会请求失败,导致后续解析报错。 - br标签提取逻辑错误:br是换行分隔标签,本身不存储文本内容,遍历br无法拿到要求文本,文本实际存储在p标签的直接子节点中。
- 变量作用域与逻辑缺陷:
respon_text += li.text中的li变量仅在ul分支的循环中定义,进入p分支时会触发未定义报错;同时ul分支仅打印内容未做存储,会导致要求字段缺失,列表长度和其他字段不匹配。
修复方案
无需单独适配ul、p等不同标签结构,直接提取要求容器的全部可见文本即可,兼容性更高,修改后的代码如下:
from os import execle, link, unlink, write from typing import Text import requests from bs4 import BeautifulSoup import csv from itertools import zip_longest job_titleL =[] company_nameL=[] location_nameL=[] experience_inL=[] links=[] salary=[] job_requirementsL=[] date=[] result= requests.get(f"https://wuzzuf.net/search/jobs/?a=%7B%7D&q=python&start=1") source = result.content soup= BeautifulSoup(source , "lxml") job_titles = soup.find_all("h2",{"class":"css-m604qf"} ) companies_names = soup.find_all("a",{"class":"css-17s97q8"}) locations_names = soup.find_all("span",{"class":"css-5wys0k"}) experience_in = soup.find_all("a", {"class":"css-5x9pm1"}) posted_new = soup.find_all("div",{"class":"css-4c4ojb"}) posted_old = soup.find_all("div",{"class":"css-do6t5g"}) posted = [*posted_new,*posted_old] for L in range(len(job_titles)): job_titleL.append(job_titles[L].text) # 补全链接域名前缀 links.append("https://wuzzuf.net" + job_titles[L].find('a').attrs['href']) company_nameL.append(companies_names[L].text) location_nameL.append(locations_names[L].text) experience_inL.append(experience_in[L].text) date_text=posted[L].text.replace("-","").strip() date.append(posted[L].text) for link in links: result= requests.get(link) source= result.content soup=BeautifulSoup(source,"lxml") # 先判断外层要求容器是否存在 requirement_container = soup.find("div",{"class":"css-1t5f0fr"}) respon_text = "" if requirement_container: # 直接提取容器内所有文本,用|替换换行作为分隔符 respon_text = "|".join([text.strip() for text in requirement_container.stripped_strings]) # 不管有没有内容都追加,保证列表长度一致 job_requirementsL.append(respon_text) file_list=[job_titleL,company_nameL,date,location_nameL,experience_inL,links,job_requirementsL] exported=zip_longest(*file_list) with open('newspeard2.csv',"w", encoding="utf-8-sig") as spreadsheet: wr=csv.writer(spreadsheet) wr.writerow(["job title", "company name","date", "location", "experience in","links","job requirements"]) wr.writerows(exported)
额外说明:新增的encoding="utf-8-sig"参数是为了保证CSV打开时特殊字符不会乱码,不需要可以删掉。
内容的提问来源于stack exchange,提问作者kamalMKA
相关产品推荐
相关产品推荐

