You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取LinkedIn职位发布时仅获取前7个href的技术问题求助

问题分析与解决方案

代码本身的错误

你的代码存在两个关键问题:

  • 提前调用未定义的cell变量,无法正确遍历每个职位链接
  • 仅解析了初始页面的静态内容,未处理LinkedIn职位列表的动态加载机制(默认只加载前7-10条,剩余内容需滚动页面触发加载)

修正步骤与代码

1. 修复基础逻辑错误

先修正选择器和循环逻辑,确保能正确提取每个职位的href:

base_url = 'https://www.linkedin.com'

# 先处理动态加载(见步骤2),再获取页面源码
page_source = driver.page_source
soup = bs(page_source, 'html.parser')

# 定位职位卡片外层div,再从中提取内部的a标签
job_cards = soup.find_all('div', attrs={'class':'full-width artdeco-entity-lockup__title ember-view'})
for card in job_cards:
    a_tag = card.find('a')
    # 确保a标签存在且包含href属性
    if a_tag and 'href' in a_tag.attrs:
        job_url = base_url + a_tag['href']
        print(job_url)

2. 处理动态加载(核心解决仅获取前7条的问题)

LinkedIn职位列表采用滚动加载,需要用Selenium模拟滚动触发更多内容加载,循环直到无新内容可加载:

import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup as bs

driver = webdriver.Chrome()
driver.get('https://www.linkedin.com/jobs/search/?currentJobId=3702922458')

# 模拟滚动加载全部职位
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待页面加载(根据网络情况调整时长)
    time.sleep(3)
    # 对比页面高度,判断是否加载新内容
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        # 尝试点击"查看更多"按钮(若页面存在该按钮)
        try:
            see_more_btn = driver.find_element(By.CSS_SELECTOR, "button.infinite-scroller__show-more-button")
            see_more_btn.click()
            time.sleep(3)
            new_height = driver.execute_script("return document.body.scrollHeight")
        except:
            # 无更多内容,退出循环
            break
    last_height = new_height

# 执行解析逻辑
page_source = driver.page_source
soup = bs(page_source, 'html.parser')

job_cards = soup.find_all('div', attrs={'class':'full-width artdeco-entity-lockup__title ember-view'})
for card in job_cards:
    a_tag = card.find('a')
    if a_tag and 'href' in a_tag.attrs:
        job_url = base_url + a_tag['href']
        print(job_url)

driver.quit()

3. 额外注意事项

  • LinkedIn有反爬机制,频繁请求可能被限制,建议延长等待时间,或搭配代理IP使用
  • 确保已登录LinkedIn(部分职位列表需要登录才能完整加载)
  • 定期检查页面元素的class属性,LinkedIn可能更新页面结构导致选择器失效

内容的提问来源于stack exchange,提问作者Anas Lachhab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 06:55:40