You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium和BeautifulSoup爬取LinkedIn时HTML标签变化,无法获取经历与教育信息

解决LinkedIn个人主页经历/教育板块爬取问题

核心问题分析

你遇到的是LinkedIn动态渲染HTML导致的定位失效:

  • 经历/教育板块的section标签ID带有随机动态数字,固定ID选择器完全失效
  • 直接依赖标签层级(比如find_all("p")[1])的解析逻辑,一旦页面结构微调就会出错

解决方案

1. 用稳定属性定位板块

放弃固定ID,改用LinkedIn页面中更稳定的数据属性或语义化标识:

  • 经历板块可通过data-section="experience"属性定位
  • 教育板块可通过data-section="education"属性定位
  • 也可通过板块标题文本(比如"Experience")反向定位父容器

2. 确保页面完全加载(Selenium部分)

LinkedIn内容是动态加载的,必须等待目标板块出现后再解析:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待经历板块加载完成(最多10秒)
wait = WebDriverWait(driver, 10)
experience_section = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'section[data-section="experience"]')))

# 获取页面源码交给BeautifulSoup解析
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'html.parser')

3. 鲁棒的经历内容提取逻辑

避免依赖固定标签索引,改用类名选择器或CSS选择器精准定位元素:

# 定位经历板块的列表容器
experience_container = soup.select_one('section[data-section="experience"] ul')
if not experience_container:
    print("未找到经历板块")
    exit()

# 遍历所有经历条目
for item in experience_container.select('li > div'):
    # 提取职位名称
    job_title = item.select_one('h3').get_text(strip=True) if item.select_one('h3') else "无职位信息"
    # 提取公司名称
    company_name = item.select_one('p[class*="company-name"]').get_text(strip=True) if item.select_one('p[class*="company-name"]') else "无公司信息"
    # 提取日期和时长
    date_info = item.select('h4 span[class*="date"]')
    joining_date = date_info[0].get_text(strip=True) if len(date_info) >=1 else ""
    employment_duration = date_info[1].get_text(strip=True) if len(date_info)>=2 else ""
    
    print(f"职位: {job_title}")
    print(f"公司: {company_name}")
    print(f"时间: {joining_date}, {employment_duration}\n")

4. 教育板块的类似处理

# 定位教育板块
education_container = soup.select_one('section[data-section="education"] ul')
if education_container:
    for item in education_container.select('li > div'):
        school_name = item.select_one('h3').get_text(strip=True) if item.select_one('h3') else "无学校信息"
        degree = item.select_one('span[class*="degree"]').get_text(strip=True) if item.select_one('span[class*="degree"]') else "无学位信息"
        field_of_study = item.select_one('span[class*="field-of-study"]').get_text(strip=True) if item.select_one('span[class*="field-of-study"]') else "无专业信息"
        graduation_date = item.select_one('span[class*="date"]').get_text(strip=True) if item.select_one('span[class*="date"]') else "无毕业日期"
        
        print(f"学校: {school_name}")
        print(f"学位/专业: {degree} {field_of_study}")
        print(f"毕业日期: {graduation_date}\n")

注意事项

  • LinkedIn有反爬机制,频繁请求可能触发验证码或账号限制,建议添加合理等待时间,避免短时间内多次访问
  • 类名可能随LinkedIn更新变化,若失效,重新查看页面元素的最新类名或数据属性,调整选择器即可

内容的提问来源于stack exchange,提问作者Nurul Atirah Binti Ahmad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 19:07:57