You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium和Python爬取LinkedIn用户教育页面嵌套数据?

解决LinkedIn教育经历爬取的Selenium脚本

核心问题分析

你遇到的no such element错误,主要原因是页面未完全加载就执行定位操作,或是LinkedIn的内容为动态异步渲染,直接定位会找不到元素。必须等待元素加载完成,并处理懒加载的滚动逻辑。

完整爬取脚本

以下是可直接运行的脚本,包含登录、滚动加载、循环提取机构名称和就读年份的逻辑:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

def scrape_linkedin_education(url):
    # 初始化Chrome浏览器(需确保chromedriver版本与浏览器匹配)
    driver = webdriver.Chrome()
    driver.get("https://www.linkedin.com/login")
    
    # 登录操作(替换为你的LinkedIn账号密码)
    username = driver.find_element(By.ID, "username")
    password = driver.find_element(By.ID, "password")
    username.send_keys("你的账号")
    password.send_keys("你的密码")
    driver.find_element(By.XPATH, "//button[@type='submit']").click()
    
    # 等待登录跳转,再打开目标教育页面
    time.sleep(3)
    driver.get(url)
    
    # 滚动加载所有教育经历(LinkedIn采用懒加载,需滚动到底部触发加载)
    last_height = driver.execute_script("return document.body.scrollHeight")
    while True:
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(2)
        new_height = driver.execute_script("return document.body.scrollHeight")
        if new_height == last_height:
            break
        last_height = new_height
    
    # 等待教育经历条目容器加载完成
    wait = WebDriverWait(driver, 10)
    education_items = wait.until(EC.presence_of_all_elements_located(
        (By.XPATH, "//div[contains(@class, 'pvs-list__item--line-separated')]")
    ))
    
    # 循环提取每个教育经历的机构名称和年份
    education_data = []
    for item in education_items:
        try:
            # 提取机构名称(匹配目标class)
            institution = item.find_element(By.XPATH, ".//span[contains(@class, 'mr1 hoverable-link-text t-bold')]").text
            # 提取就读年份(匹配年份所在的灰色文本class)
            year = item.find_element(By.XPATH, ".//span[contains(@class, 't-black--light')]").text
            education_data.append({
                "机构名称": institution,
                "就读年份": year
            })
        except Exception as e:
            # 跳过解析失败的条目(避免单个条目错误中断整个程序)
            print(f"解析条目失败: {e}")
            continue
    
    # 关闭浏览器
    driver.quit()
    return education_data

# 测试脚本(替换为目标用户的教育页面链接)
if __name__ == "__main__":
    target_url = "https://www.linkedin.com/in/williamhgates/details/education/"
    result = scrape_linkedin_education(target_url)
    for idx, edu in enumerate(result, 1):
        print(f"第{idx}条教育经历:")
        print(f"机构: {edu['机构名称']}")
        print(f"年份: {edu['就读年份']}\n")

关键注意事项

  • chromedriver版本:必须与你的Chrome浏览器版本完全匹配,否则无法启动浏览器。
  • 登录限制:LinkedIn会检测自动化行为,频繁登录可能触发验证码,建议手动登录后注释掉登录代码,直接打开目标页面。
  • 元素定位容错:LinkedIn的class名称可能会微调,若后续定位失败,可通过浏览器开发者工具重新获取最新的元素定位表达式。
  • 反爬规避:避免过快滚动或请求,适当增加等待时间,防止账号被LinkedIn封禁。

内容的提问来源于stack exchange,提问作者Nadia S.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 07:24:24