如何用BeautifulSoup提取网页嵌套类结构的招聘数据?
解决BeautifulSoup提取招聘信息返回None的问题
你的代码返回None的核心原因是:你直接查找的是<i>标签(比如class='i-globe'),但这类标签本身没有文本内容,实际的地点、公司等信息是在<i>标签的相邻文本节点中,或者说在它的父元素<span class="info">里。以下是两种可行的解决方法:
方法一:通过<i>标签定位后取相邻文本
利用next_sibling获取<i>标签后面的文本节点,再用strip()清理多余空格:
import requests from bs4 import BeautifulSoup URL = "https://pythonjobs.github.io/" page = requests.get(URL) soup = BeautifulSoup(page.content, 'html.parser') job_container = soup.find(id='container') job_elements = job_container.find_all(class_='job') for job_element in job_elements: # 提取职位名称 title = job_element.find('h1').get_text(strip=True) # 提取地点 location_i = job_element.find('i', class_='i-globe') location = location_i.next_sibling.strip() if location_i else "未知地点" # 提取发布日期 date_i = job_element.find('i', class_='i-calendar') date = date_i.next_sibling.strip() if date_i else "未知日期" # 提取工作类型 length_i = job_element.find('i', class_='i-chair') length = length_i.next_sibling.strip() if length_i else "未知工作类型" # 提取公司名称 company_i = job_element.find('i', class_='i-company') company = company_i.next_sibling.strip() if company_i else "未知公司" # 提取职位描述 description_elem = job_element.find('p', class_='detail') description = description_elem.get_text(strip=True) if description_elem else "无描述" # 打印结果 print(f"职位: {title}") print(f"地点: {location}") print(f"发布日期: {date}") print(f"工作类型: {length}") print(f"公司: {company}") print(f"描述: {description}\n")
方法二:遍历所有<span class="info">元素分类提取
先获取所有info类型的span,再根据内部<i>标签的class来区分信息类型:
import requests from bs4 import BeautifulSoup URL = "https://pythonjobs.github.io/" page = requests.get(URL) soup = BeautifulSoup(page.content, 'html.parser') job_container = soup.find(id='container') job_elements = job_container.find_all(class_='job') for job_element in job_elements: title = job_element.find('h1').get_text(strip=True) info_spans = job_element.find_all('span', class_='info') # 初始化默认值 location = date = length = company = "未知" for span in info_spans: i_tag = span.find('i') if not i_tag: continue # 根据i标签的class判断信息类型 if 'i-globe' in i_tag.get('class'): location = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip() elif 'i-calendar' in i_tag.get('class'): date = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip() elif 'i-chair' in i_tag.get('class'): length = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip() elif 'i-company' in i_tag.get('class'): company = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip() description_elem = job_element.find('p', class_='detail') description = description_elem.get_text(strip=True) if description_elem else "无描述" # 打印结果 print(f"职位: {title}") print(f"地点: {location}") print(f"发布日期: {date}") print(f"工作类型: {length}") print(f"公司: {company}") print(f"描述: {description}\n")
注意事项
- 两种方法都做了空值判断,避免因页面结构变化导致的报错
get_text(strip=True)用于清理文本中的多余换行和空格- 方法一更直接高效,方法二更灵活,适合页面info顺序可能变化的场景
内容的提问来源于stack exchange,提问作者Ruibin Wu
相关产品推荐
相关产品推荐

