解析div触发AttributeError:'NoneType'对象无'find'属性求助
问题:爬取Fake Jobs页面时触发AttributeError错误
问题详情
爬取Fake Jobs站点职位信息时,解析职位详情页内容出现以下错误:
line 38, in <module> item_element_elements = results_site.find("div", class_="content") AttributeError: 'NoneType' object has no attribute 'find'
出错核心代码片段:
item_site = requests.get(item_element["href"]) item_soup = BeautifulSoup(item_site.content, "html.parser") results_site = item_soup.find(id="ResultsContainer") item_element_elements = results_site.find("div", class_="content") item_element_element = item_element_elements.find("p", class_=False)
完整爬虫代码:
import requests from bs4 import BeautifulSoup from texttable import Texttable url = "https://realpython.github.io/fake-jobs/" site = requests.get(url) #send a request to the site table = Texttable() #create a table table.set_chars(['-', '|', '+', '=']) table.header(['Titel','Company','Location']) table.set_cols_dtype(['t','i','a']) table.set_cols_align(["c", "c", "c"]) table.set_cols_valign(["m", "m", "m"]) table.set_cols_width([20,20,20]) table.set_deco(Texttable.BORDER|Texttable.HEADER |Texttable.HLINES| Texttable.VLINES) with open('Shore.txt', 'w') as f: #create a file pass soup = BeautifulSoup(site.content, "html.parser") results = soup.find(id="ResultsContainer") job_elements = results.find_all("div", class_="card-content") #find all div with class "card-content" for job_element in job_elements: title_element = job_element.find("h2", class_="title") #get the different elements from divs with class "card-content" company_element = job_element.find("h3", class_="company") #get the different elements from divs with class "card-content" location_element = job_element.find("p", class_="location") #get the different elements from divs with class "card-content" item_element = job_element.find("a", class_="card-footer-item") #get the link with divs from class "card-content" item_site = requests.get(item_element["href"]) #send a request to the site from link item_soup = BeautifulSoup(item_site.content, "html.parser") results_site = item_soup.find(id="ResultsContainer") item_element_elements = results_site.find("div", class_="content") item_element_element = item_element_elements.find("p", class_=False) #get the text without class print(title_element.text.strip()) #get it all data received into the console print(company_element.text.strip()) #get all data received into the console print(location_element.text.strip()) #get all data received into the console print(item_element_element.text.strip()) # get all data received into the console table.add_row([title_element.text.strip(),company_element.text.strip(),location_element.text.strip()]) #add rows in corrects rows "add_rows" with open('Shore.txt', 'w') as f: #enter all data received into a table file f.write(table.draw()) f.write(str(len(job_elements))) f.close print(len(job_elements)) #get the number of elements with the class
错误原因
职位列表页存在id="ResultsContainer"的容器,但每个职位的详情页结构和列表页完全不同,详情页中没有id为ResultsContainer的元素,导致results_site = item_soup.find(id="ResultsContainer")返回None,后续调用None.find()直接触发AttributeError。
解决方案
1. 修正详情页元素定位逻辑
直接跳过ResultsContainer的查找,详情页内容直接在div.content中,修改代码如下:
item_site = requests.get(item_element["href"]) item_soup = BeautifulSoup(item_site.content, "html.parser") # 直接查找content容器,无需找ResultsContainer item_element_elements = item_soup.find("div", class_="content") if item_element_elements: # 先判断元素是否存在,避免报错 item_element_element = item_element_elements.find("p", class_=False) if item_element_element: print(item_element_element.text.strip())
2. 增加异常与边界处理
为避免网络请求失败、元素缺失导致程序崩溃,补充以下处理:
- 检查请求详情页的响应状态码
- 对每个
find操作的结果做非空判断 - 捕获网络请求异常
3. 优化文件写入逻辑
原代码每次循环都以w模式打开文件,会覆盖之前的内容,改为循环结束后统一写入:
# 循环结束后统一写入文件 with open('Shore.txt', 'w') as f: f.write(table.draw()) f.write(f"\nTotal jobs: {len(job_elements)}")
修改后的完整代码
import requests from bs4 import BeautifulSoup from texttable import Texttable url = "https://realpython.github.io/fake-jobs/" site = requests.get(url) table = Texttable() table.set_chars(['-', '|', '+', '=']) table.header(['Title','Company','Location']) table.set_cols_dtype(['t','t','t']) # 修正数据类型:Company是文本不是整数 table.set_cols_align(["c", "c", "c"]) table.set_cols_valign(["m", "m", "m"]) table.set_cols_width([20,20,20]) table.set_deco(Texttable.BORDER|Texttable.HEADER |Texttable.HLINES| Texttable.VLINES) soup = BeautifulSoup(site.content, "html.parser") results = soup.find(id="ResultsContainer") job_elements = results.find_all("div", class_="card-content") for job_element in job_elements: title_element = job_element.find("h2", class_="title") company_element = job_element.find("h3", class_="company") location_element = job_element.find("p", class_="location") item_element = job_element.find("a", class_="card-footer-item") # 提取基础信息 title = title_element.text.strip() company = company_element.text.strip() location = location_element.text.strip() print(title, company, location, sep="\n") # 处理详情页 if item_element and "href" in item_element.attrs: item_url = item_element["href"] try: item_site = requests.get(item_url) item_site.raise_for_status() # 抛出HTTP错误 item_soup = BeautifulSoup(item_site.content, "html.parser") content_div = item_soup.find("div", class_="content") if content_div: detail_text = content_div.find("p", class_=False) if detail_text: print(detail_text.text.strip()) print("---") except requests.exceptions.RequestException as e: print(f"获取详情页失败: {e}") print("---") table.add_row([title, company, location]) # 统一写入文件 with open('Shore.txt', 'w') as f: f.write(table.draw()) f.write(f"\nTotal jobs: {len(job_elements)}") print(f"共找到职位: {len(job_elements)}个")
额外修正点
- 修正
table.set_cols_dtype(['t','i','a']):Company是文本类型,改为t避免表格渲染报错 - 增加HTTP请求异常捕获,处理网络请求失败场景
- 对元素属性做存在性判断,避免
item_element["href"]触发KeyError
内容的提问来源于stack exchange,提问作者MrYarikPasik
相关产品推荐
相关产品推荐

