如何用Selenium+Python定位LinkedIn经历板块并提取简历信息?
解决LinkedIn简历经历&教育板块的Selenium定位与信息提取问题
核心问题原因
你的定位失败主要是LinkedIn页面结构频繁更新,旧教程/问答里的选择器已经失效;另外未登录LinkedIn的话,简历内容会被限制,也会导致元素找不到或超时。
前置要求
- 必须先登录LinkedIn(可以用Selenium自动登录,或手动登录后复用浏览器会话,避免验证码问题)
- 访问目标简历页面后,需要滚动到目标板块区域,触发内容加载
正确的定位与提取代码示例
以示例简历为例,以下是可直接复用的Python+Selenium代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # 初始化Chrome浏览器(可替换为其他浏览器) driver = webdriver.Chrome() wait = WebDriverWait(driver, 15) # 1. 登录LinkedIn(建议手动登录,自动登录易触发验证码) driver.get("https://www.linkedin.com/login") input("登录完成后按回车继续...") # 2. 访问目标简历页面 driver.get("https://www.linkedin.com/in/kendra-tyson/") time.sleep(2) # 等待页面初步加载 # 3. 滚动到经历板块,触发懒加载 experience_section = wait.until(EC.presence_of_element_located((By.ID, "experience"))) driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", experience_section) time.sleep(1) # 4. 定位所有经历条目 experience_items = wait.until(EC.visibility_of_all_elements_located( (By.CSS_SELECTOR, "section#experience li.pvs-list__item--line-separated") )) # 5. 提取每个经历的详细信息 print("=== 经历板块信息 ===") for item in experience_items: # 职位名称 job_title = item.find_element(By.CSS_SELECTOR, "span[aria-hidden='true']").text.strip() # 公司名称 company = item.find_element(By.CSS_SELECTOR, "span.t-14.t-normal").text.strip() # 时间线与全职状态 time_status = item.find_element(By.CSS_SELECTOR, "span.t-14.t-normal.t-black--light").text.strip() # 地点(处理无地点的情况) location_elems = item.find_elements(By.CSS_SELECTOR, "span.t-14.t-normal.t-black--light") location = location_elems[2].text.strip() if len(location_elems) >=3 else "无地点信息" print(f"职位: {job_title}") print(f"公司: {company}") print(f"时间&状态: {time_status}") print(f"地点: {location}") print("---") # 6. 教育板块提取(复用类似逻辑) print("\n=== 教育板块信息 ===") education_section = wait.until(EC.presence_of_element_located((By.ID, "education"))) driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", education_section) time.sleep(1) education_items = wait.until(EC.visibility_of_all_elements_located( (By.CSS_SELECTOR, "section#education li.pvs-list__item--line-separated") )) for item in education_items: school = item.find_element(By.CSS_SELECTOR, "span[aria-hidden='true']").text.strip() degree_major = item.find_element(By.CSS_SELECTOR, "span.t-14.t-normal").text.strip() time_period = item.find_element(By.CSS_SELECTOR, "span.t-14.t-normal.t-black--light").text.strip() print(f"院校: {school}") print(f"学位&专业: {degree_major}") print(f"就读时间: {time_period}") print("---") driver.quit()
关键说明
- 用
section#experience定位板块容器,当前LinkedIn的经历板块是section标签,比旧的div定位更准确 - 每个经历条目用
li.pvs-list__item--line-separated,这是2024年有效的class,替代了你之前使用的过时选择器 - 滚动到板块区域是必要操作,LinkedIn采用懒加载机制,不滚动的话部分元素不会渲染
- 若遇到验证码,建议手动登录后复用浏览器会话(例如用Chrome用户数据目录启动),降低反爬触发概率
内容的提问来源于stack exchange,提问作者Jesper Ezra
相关产品推荐
相关产品推荐

