You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取网页嵌套类结构的招聘数据?

解决BeautifulSoup提取招聘信息返回None的问题

你的代码返回None的核心原因是:你直接查找的是<i>标签(比如class='i-globe'),但这类标签本身没有文本内容,实际的地点、公司等信息是在<i>标签的相邻文本节点中,或者说在它的父元素<span class="info">里。以下是两种可行的解决方法:

方法一:通过<i>标签定位后取相邻文本

利用next_sibling获取<i>标签后面的文本节点,再用strip()清理多余空格:

import requests
from bs4 import BeautifulSoup

URL = "https://pythonjobs.github.io/"
page = requests.get(URL)

soup = BeautifulSoup(page.content, 'html.parser')

job_container = soup.find(id='container')
job_elements = job_container.find_all(class_='job')

for job_element in job_elements:
    # 提取职位名称
    title = job_element.find('h1').get_text(strip=True)
    
    # 提取地点
    location_i = job_element.find('i', class_='i-globe')
    location = location_i.next_sibling.strip() if location_i else "未知地点"
    
    # 提取发布日期
    date_i = job_element.find('i', class_='i-calendar')
    date = date_i.next_sibling.strip() if date_i else "未知日期"
    
    # 提取工作类型
    length_i = job_element.find('i', class_='i-chair')
    length = length_i.next_sibling.strip() if length_i else "未知工作类型"
    
    # 提取公司名称
    company_i = job_element.find('i', class_='i-company')
    company = company_i.next_sibling.strip() if company_i else "未知公司"
    
    # 提取职位描述
    description_elem = job_element.find('p', class_='detail')
    description = description_elem.get_text(strip=True) if description_elem else "无描述"
    
    # 打印结果
    print(f"职位: {title}")
    print(f"地点: {location}")
    print(f"发布日期: {date}")
    print(f"工作类型: {length}")
    print(f"公司: {company}")
    print(f"描述: {description}\n")

方法二:遍历所有<span class="info">元素分类提取

先获取所有info类型的span,再根据内部<i>标签的class来区分信息类型:

import requests
from bs4 import BeautifulSoup

URL = "https://pythonjobs.github.io/"
page = requests.get(URL)

soup = BeautifulSoup(page.content, 'html.parser')

job_container = soup.find(id='container')
job_elements = job_container.find_all(class_='job')

for job_element in job_elements:
    title = job_element.find('h1').get_text(strip=True)
    info_spans = job_element.find_all('span', class_='info')
    
    # 初始化默认值
    location = date = length = company = "未知"
    
    for span in info_spans:
        i_tag = span.find('i')
        if not i_tag:
            continue
        
        # 根据i标签的class判断信息类型
        if 'i-globe' in i_tag.get('class'):
            location = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip()
        elif 'i-calendar' in i_tag.get('class'):
            date = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip()
        elif 'i-chair' in i_tag.get('class'):
            length = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip()
        elif 'i-company' in i_tag.get('class'):
            company = span.get_text(strip=True).replace(i_tag.get_text(strip=True), '').strip()
    
    description_elem = job_element.find('p', class_='detail')
    description = description_elem.get_text(strip=True) if description_elem else "无描述"
    
    # 打印结果
    print(f"职位: {title}")
    print(f"地点: {location}")
    print(f"发布日期: {date}")
    print(f"工作类型: {length}")
    print(f"公司: {company}")
    print(f"描述: {description}\n")

注意事项

  • 两种方法都做了空值判断,避免因页面结构变化导致的报错
  • get_text(strip=True)用于清理文本中的多余换行和空格
  • 方法一更直接高效,方法二更灵活,适合页面info顺序可能变化的场景

内容的提问来源于stack exchange,提问作者Ruibin Wu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 12:46:58