You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python无痕模式爬取LinkedIn全部职位数据(解决仅获25条问题)

问题分析与解决方案

核心问题

你的代码存在两个关键问题导致只能获取25条数据:

  1. 误用请求工具:同时使用requests.get()和Selenium浏览器实例,但requests获取的是LinkedIn静态初始页面,完全没利用Selenium的动态渲染能力,自然拿不到后续加载的职位数据。
  2. 缺少动态加载逻辑:LinkedIn职位列表采用「滚动加载+点击Load More」的方式加载更多内容,你的代码没有实现这部分交互逻辑。

解决步骤

  1. 移除requests依赖:直接用Selenium驱动浏览器访问页面,等待渲染完成后获取页面源码。
  2. 实现滚动加载:循环滚动至页面底部,每次滚动后等待内容加载。
  3. 处理Load More按钮:当滚动触发按钮出现时,定位并点击,直到按钮不再显示(无更多数据)。
  4. 添加显式等待:用Selenium的显式等待确保元素加载完成,避免因页面未加载完全导致的定位失败。

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import ElementNotInteractableException, TimeoutException

from bs4 import BeautifulSoup as beauty
import time

chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("incognito")
# 规避LinkedIn自动化检测
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)

url_link = 'https://www.linkedin.com/jobs/search/?currentJobId=3187861296&geoId=102713980&keywords=mckinsey&location=India&refresh=true'

driver = webdriver.Chrome(
    service=Service(ChromeDriverManager().install()), options=chrome_options
)
# 隐藏webdriver标识
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

try:
    driver.get(url_link)
    print(f"{url_link} 已加载,开始爬取职位数据")

    # 等待初始职位列表加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "job-search-card"))
    )

    last_height = driver.execute_script("return document.body.scrollHeight")
    load_more_exists = True

    while load_more_exists:
        # 滚动至页面底部
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # 等待内容加载
        time.sleep(3)

        # 尝试点击Load More按钮
        try:
            load_more_btn = WebDriverWait(driver, 5).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, "button.infinite-scroller__show-more-button"))
            )
            load_more_btn.click()
            time.sleep(3)
        except (ElementNotInteractableException, TimeoutException):
            # 按钮不存在或不可点击,检查页面高度是否变化
            new_height = driver.execute_script("return document.body.scrollHeight")
            if new_height == last_height:
                load_more_exists = False
            last_height = new_height

    # 获取完整渲染后的页面源码
    page_source = driver.page_source
    soup = beauty(page_source, "html.parser")

    jobs = soup.find_all(
        "div",
        class_="base-card relative w-full hover:no-underline focus:no-underline base-card--link base-search-card base-search-card--link job-search-card",
    )

    # 将所有职位信息写入单个文件,便于管理
    with open("linkedin_jobs.txt", "w", encoding="utf-8") as f:
        for idx, job in enumerate(jobs, 1):
            try:
                job_title = job.find("h3", class_="base-search-card__title").text.strip()
                job_company = job.find("h4", class_="base-search-card__subtitle").text.strip()
                job_location = job.find("span", class_="job-search-card__location").text.strip()
                job_link = job.find("a", class_="base-card__full-link")["href"]

                job_info = f"职位{idx}:\n标题: {job_title}\n公司: {job_company}\n地点: {job_location}\n链接: {job_link}\n---\n"
                print(job_info)
                f.write(job_info)
            except Exception as e:
                print(f"处理职位时出错: {e}")
                continue

finally:
    driver.quit()

关键说明

  • 反爬优化:添加了禁用自动化检测的参数,避免LinkedIn识别出Selenium爬虫。
  • 动态加载逻辑:循环滚动页面并处理Load More按钮,直到页面高度不再变化(无更多数据)。
  • 可靠性提升:用WebDriverWait替代固定休眠,确保元素加载完成后再进行操作。
  • 文件管理优化:将所有职位信息合并写入一个文件,避免多文件分散的问题。

内容的提问来源于stack exchange,提问作者victor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:42:30