You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium打开LinkedIn主页提取帖子点赞者姓名与职位?

用Selenium提取LinkedIn帖子点赞用户姓名的解决方案

原代码存在的问题

  • 直接使用find_element_by_class_name会在元素不存在时抛出异常,导致循环直接中断
  • 固定time.sleep(5)不够灵活,页面加载速度不稳定时,要么浪费时间要么加载不充分
  • 仅靠时间阈值判断停止,无法精准识别是否已加载完所有点赞者

核心改进方案

  • 用显式等待替代固定sleep,等待元素可点击或存在,提升效率和稳定性
  • 加入异常捕获逻辑,当"显示更多"按钮不存在或无法点击时,自动终止加载循环
  • 分阶段提取已加载的用户信息,避免重复处理
  • 处理页面跳转后的返回操作,确保能持续提取后续用户数据

完整实现代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException
import time

# 初始化Chrome浏览器(需提前配置对应驱动)
driver = webdriver.Chrome()
driver.get("你的LinkedIn帖子页面URL")

# 建议手动完成登录(LinkedIn反爬严格,自动登录易触发限制)
wait = WebDriverWait(driver, 10)

# 打开点赞弹窗
like_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(@aria-label, '点赞')]")))
like_button.click()

# 定位点赞弹窗容器
likes_container = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "artdeco-modal__content")))

start_time = time.time()
max_wait_time = 120  # 最大等待时长,防止无限循环

# 循环加载所有点赞者
while True:
    try:
        # 等待"显示更多"按钮可点击
        show_more_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "display-flex.p5")))
        # 滚动到按钮位置,避免被页面元素遮挡
        driver.execute_script("arguments[0].scrollIntoView();", show_more_btn)
        show_more_btn.click()
        print("已加载更多点赞者")
        time.sleep(2)  # 给页面加载留缓冲时间
    except (NoSuchElementException, ElementClickInterceptedException):
        # 无更多按钮或无法点击,说明加载完成
        print("已加载所有点赞者")
        break
    except Exception as e:
        print(f"加载出错: {str(e)}")
        time.sleep(3)
    
    # 超过最大等待时长则终止
    if time.time() - start_time > max_wait_time:
        print("超过最大等待时间,终止加载")
        break

# 提取已加载的用户基础信息
user_items = likes_container.find_elements(By.CLASS_NAME, "artdeco-list__item")
user_info_list = []

for item in user_items:
    try:
        name = item.find_element(By.CLASS_NAME, "artdeco-entity-lockup__title").text.strip()
        profile_link = item.find_element(By.TAG_NAME, "a").get_attribute("href")
        user_info_list.append({"姓名": name, "主页链接": profile_link})
        print(f"已提取用户: {name}")
    except Exception as e:
        print(f"提取用户信息出错: {str(e)}")
        continue

# 逐个打开主页获取更多信息(示例:提取职位)
for user in user_info_list:
    try:
        driver.get(user["主页链接"])
        wait.until(EC.presence_of_element_located((By.CLASS_NAME, "text-body-medium")))
        position = driver.find_element(By.CLASS_NAME, "text-body-medium").text.strip()
        user["职位"] = position
        print(f"{user['姓名']} 的职位: {position}")
        # 返回点赞弹窗页面,继续处理下一个用户
        driver.back()
        wait.until(EC.presence_of_element_located((By.CLASS_NAME, "artdeco-modal__content")))
        time.sleep(1)
    except Exception as e:
        print(f"获取{user['姓名']}主页信息出错: {str(e)}")
        driver.back()
        continue

# 关闭浏览器
driver.quit()

# 打印最终结果
print("\n提取的用户信息:")
for info in user_info_list:
    print(info)

关键注意事项

  • 反爬规避:LinkedIn对自动化操作检测严格,建议在操作中加入随机等待时间,避免短时间内大量请求触发验证码或账号限制
  • 元素更新:LinkedIn的页面元素class名称可能会不定期更新,若代码失效,需重新检查元素的定位符(class/xpath)
  • 登录建议:优先手动完成登录后再运行代码,自动登录逻辑容易被平台检测到,导致账号受限

内容的提问来源于stack exchange,提问作者Ashish Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 02:06:03