You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup通过div class_=css-gz8dae获取职位描述?求可用代码

解决JustJoin.it职位描述爬虫问题

问题说明

作为Python爬虫新手,尝试用BeautifulSoup抓取justjoin.it职位页面(示例链接:https://justjoin.it/offers/itds-net-fullstack-developer-angular)的职位描述,但原有代码无法正常工作——此前更换div类名在其他招聘网站可成功获取内容。原有代码如下:

import requests
from bs4 import BeautifulSoup

link="https://justjoin.it/offers/jungle-devops-engineer"

response_IDs=requests.get(link)
soup=BeautifulSoup(response_IDs.text, 'html.parser')
Search_part = soup.find(id='root')
description= Search_part.find_all('div', class_='css-gz8dae')

for i in description:
    print(i)

问题原因

justjoin.it的页面采用动态渲染,且类名(如css-gz8dae)是前端框架自动生成的随机类名,会随页面部署变化;直接用requests.get获取的源码中,职位描述内容尚未被前端渲染出来,导致无法通过固定类名定位元素。

解决方案

方法1:调用官方API(推荐)

justjoin.it的职位数据通过公开API接口返回,直接请求API可高效获取结构化数据,无需处理页面渲染问题:

import requests
from bs4 import BeautifulSoup

# 替换为目标职位链接的最后部分(即offer的slug)
offer_slug = "jungle-devops-engineer"
api_url = f"https://justjoin.it/api/offers/{offer_slug}"

# 发送API请求
response = requests.get(api_url)
response.raise_for_status()  # 捕获请求异常
offer_data = response.json()

# 获取原始HTML格式的职位描述
job_description_html = offer_data["description"]
print("HTML格式描述:")
print(job_description_html)

# 若需提取纯文本,用BeautifulSoup解析
soup = BeautifulSoup(job_description_html, "html.parser")
job_description_text = soup.get_text(strip=True)
print("\n纯文本描述:")
print(job_description_text)

方法2:用Selenium渲染页面(适合必须处理渲染后页面的场景)

如果需要模拟浏览器渲染页面,可使用Selenium配合BeautifulSoup,通过稳定的data-testid属性定位元素(比类名更可靠):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

link = "https://justjoin.it/offers/jungle-devops-engineer"

# 初始化Chrome浏览器驱动(需提前下载对应版本的ChromeDriver并配置环境变量)
driver = webdriver.Chrome()
driver.get(link)

# 等待职位描述元素加载完成
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '[data-testid="offer-description"]')))

# 解析渲染后的页面源码
soup = BeautifulSoup(driver.page_source, "html.parser")
description_element = soup.select_one('[data-testid="offer-description"]')

print("职位描述纯文本:")
print(description_element.get_text(strip=True))

# 关闭浏览器
driver.quit()

内容的提问来源于stack exchange,提问作者Paweł Paprzycki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 04:57:37