You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何抓取JavaScript动态加载页面?GDIT招聘页爬取失败求解

问题原因分析
  • 该站点的招聘岗位数据属于客户端异步渲染内容:初始HTML仅返回JS框架代码,岗位数据是页面加载完成后,通过JS额外调用后端接口拉取、再动态渲染到页面上的,因此直接请求初始URL无法拿到对应内容。
  • 你之前方案的共性错误:
    • Beautiful Soup、未正确配置的ScraperAPI仅能获取初始静态HTML,自然拿不到动态渲染的h3标签
    • HTMLSession、Selenium的写法存在逻辑错误:你在Selenium中先获取page_source再执行sleep,还没等页面渲染完就已经拿了源码,当然找不到目标内容;HTMLSession的render方法默认等待时间不足,也会出现渲染不完整的问题。
可行解决方案

方案1:直接调用后端数据接口(最高效,无需渲染JS)

这是最推荐的方式,跳过页面渲染步骤,直接拉取结构化的JSON数据:

  1. 该站点的招聘搜索接口为:https://www.gdit.com/careers/wp-json/agency/v1/search,直接携带搜索参数发起GET请求即可
  2. 示例代码:
import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
params = {
    "q": "bossier city",
    "per_page": 100 # 可以调整每页返回数量,避免分页
}
response = requests.get("https://www.gdit.com/careers/wp-json/agency/v1/search", headers=headers, params=params)
job_list = response.json()["results"] # 接口返回的JSON里直接包含所有岗位结构化数据

for job in job_list:
    print(job["title"]) # 对应页面上h3标签的岗位名称,还有薪资、地点等字段可直接提取

方案2:修正Selenium写法,等待元素渲染完成

如果一定要模拟浏览器渲染,调整执行顺序,增加显式等待即可:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

options = webdriver.ChromeOptions()
options.add_argument("start-maximized")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
# 可选添加无头模式,不弹出浏览器窗口
# options.add_argument("--headless=new")

try:
    driver = webdriver.Chrome(options=options)
    driver.get("https://www.gdit.com/careers/search/?q=bossier%20city")
    # 显式等待h3元素加载,最长等待10秒,比固定sleep更可靠
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.TAG_NAME, "h3"))
    )
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "html.parser")
    for job in soup.find_all('h3'):
        print(job.get_text())
    driver.quit()
except Exception as e:
    print(f"加载失败: {str(e)}")
额外说明

不需要额外模拟用户输入关键词,你用的带q参数的URL本身就是搜索结果页,只要等渲染完成或直接调用接口就能拿到数据。

内容的提问来源于stack exchange,提问作者Brandon Jacobson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 17:27:04