如何用BeautifulSoup通过div class_=css-gz8dae获取职位描述?求可用代码
解决JustJoin.it职位描述爬虫问题
问题说明
作为Python爬虫新手,尝试用BeautifulSoup抓取justjoin.it职位页面(示例链接:https://justjoin.it/offers/itds-net-fullstack-developer-angular)的职位描述,但原有代码无法正常工作——此前更换div类名在其他招聘网站可成功获取内容。原有代码如下:
import requests from bs4 import BeautifulSoup link="https://justjoin.it/offers/jungle-devops-engineer" response_IDs=requests.get(link) soup=BeautifulSoup(response_IDs.text, 'html.parser') Search_part = soup.find(id='root') description= Search_part.find_all('div', class_='css-gz8dae') for i in description: print(i)
问题原因
justjoin.it的页面采用动态渲染,且类名(如css-gz8dae)是前端框架自动生成的随机类名,会随页面部署变化;直接用requests.get获取的源码中,职位描述内容尚未被前端渲染出来,导致无法通过固定类名定位元素。
解决方案
方法1:调用官方API(推荐)
justjoin.it的职位数据通过公开API接口返回,直接请求API可高效获取结构化数据,无需处理页面渲染问题:
import requests from bs4 import BeautifulSoup # 替换为目标职位链接的最后部分(即offer的slug) offer_slug = "jungle-devops-engineer" api_url = f"https://justjoin.it/api/offers/{offer_slug}" # 发送API请求 response = requests.get(api_url) response.raise_for_status() # 捕获请求异常 offer_data = response.json() # 获取原始HTML格式的职位描述 job_description_html = offer_data["description"] print("HTML格式描述:") print(job_description_html) # 若需提取纯文本,用BeautifulSoup解析 soup = BeautifulSoup(job_description_html, "html.parser") job_description_text = soup.get_text(strip=True) print("\n纯文本描述:") print(job_description_text)
方法2:用Selenium渲染页面(适合必须处理渲染后页面的场景)
如果需要模拟浏览器渲染页面,可使用Selenium配合BeautifulSoup,通过稳定的data-testid属性定位元素(比类名更可靠):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup link = "https://justjoin.it/offers/jungle-devops-engineer" # 初始化Chrome浏览器驱动(需提前下载对应版本的ChromeDriver并配置环境变量) driver = webdriver.Chrome() driver.get(link) # 等待职位描述元素加载完成 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '[data-testid="offer-description"]'))) # 解析渲染后的页面源码 soup = BeautifulSoup(driver.page_source, "html.parser") description_element = soup.select_one('[data-testid="offer-description"]') print("职位描述纯文本:") print(description_element.get_text(strip=True)) # 关闭浏览器 driver.quit()
内容的提问来源于stack exchange,提问作者Paweł Paprzycki
相关产品推荐
相关产品推荐

