You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的Selenium无法获取网页全部元素问题求助

解决LinkedIn职位列表仅能识别前7个的问题

你遇到的问题是LinkedIn职位列表采用动态加载机制——只有当元素滚动到可视区域内时,页面才会渲染对应的职位卡片。直接调用find_elements只能获取当前已经加载完成的前7个,未进入视口的元素还没被渲染,自然抓不到。

解决思路是模拟滚动职位列表区域,触发所有职位卡片的加载,之后再一次性获取全部元素。具体步骤如下:

  • 定位职位列表的滚动容器(而非整个页面),LinkedIn的职位列表通常在class为jobs-search-results-list的容器内
  • 循环滚动容器,每次滚动到底部后等待新内容加载
  • 重复滚动直到没有新的职位卡片加载出来,再执行元素定位

修改后的代码如下:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# initial setup #
chrome_driver_path = Service("C:\chromedriver.exe")
options = webdriver.ChromeOptions()
options.add_argument("--kiosk")
options.add_argument("--no-sandbox")

driver = webdriver.Chrome(service=chrome_driver_path, options=options)
driver.get("https://www.linkedin.com/jobs/search/?currentJobId=3271828611&keywords=python%20developer")

# Login #
WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, "Sign in"))).click()

WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.ID, "username"))).send_keys("email@address.com")
password = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.ID, "password")))
password.send_keys("password")
password.send_keys(Keys.ENTER)

# 等待职位列表加载完成
WebDriverWait(driver, 15).until(EC.presence_of_element_located((By.CLASS_NAME, "jobs-search-results-list")))

# 定位职位列表容器
job_list_container = driver.find_element(By.CLASS_NAME, "jobs-search-results-list")

# 滚动加载所有职位
previous_job_count = 0
while True:
    # 滚动到容器底部
    driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight", job_list_container)
    time.sleep(2)  # 等待加载时间可根据网络情况调整
    
    # 获取当前已加载的职位数量
    current_jobs = driver.find_elements(By.CSS_SELECTOR, ".job-card-container--clickable")
    current_job_count = len(current_jobs)
    
    # 如果数量不再增加,说明所有职位已加载
    if current_job_count == previous_job_count:
        break
    previous_job_count = current_job_count

# 现在获取全部职位
jobs = driver.find_elements(By.CSS_SELECTOR, ".job-card-container--clickable")
print(f"已获取全部职位,共{len(jobs)}个")

# 后续的保存和关注逻辑可以放在这里
# for job in jobs:
#     job.click()
#     WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CLASS_NAME, "jobs-save-button"))).click()
#     # 滚动到公司关注按钮区域
#     follow_btn = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CLASS_NAME, "follow")))
#     driver.execute_script("arguments[0].scrollIntoView();", follow_btn)
#     follow_btn.click()
#     time.sleep(1)

关键改进点:

  1. 用WebDriverWait替代time.sleep,提升代码稳定性,避免因网络延迟导致的元素定位失败
  2. 针对职位列表容器滚动,而非整个页面,更精准触发LinkedIn的动态加载逻辑
  3. 通过对比滚动前后的职位数量,确保所有职位都已加载完成

内容的提问来源于stack exchange,提问作者Ajay101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 01:01:11