如何用Selenium打开LinkedIn主页提取帖子点赞者姓名与职位?
用Selenium提取LinkedIn帖子点赞用户姓名的解决方案
原代码存在的问题
- 直接使用
find_element_by_class_name会在元素不存在时抛出异常,导致循环直接中断 - 固定
time.sleep(5)不够灵活,页面加载速度不稳定时,要么浪费时间要么加载不充分 - 仅靠时间阈值判断停止,无法精准识别是否已加载完所有点赞者
核心改进方案
- 用显式等待替代固定sleep,等待元素可点击或存在,提升效率和稳定性
- 加入异常捕获逻辑,当"显示更多"按钮不存在或无法点击时,自动终止加载循环
- 分阶段提取已加载的用户信息,避免重复处理
- 处理页面跳转后的返回操作,确保能持续提取后续用户数据
完整实现代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException import time # 初始化Chrome浏览器(需提前配置对应驱动) driver = webdriver.Chrome() driver.get("你的LinkedIn帖子页面URL") # 建议手动完成登录(LinkedIn反爬严格,自动登录易触发限制) wait = WebDriverWait(driver, 10) # 打开点赞弹窗 like_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(@aria-label, '点赞')]"))) like_button.click() # 定位点赞弹窗容器 likes_container = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "artdeco-modal__content"))) start_time = time.time() max_wait_time = 120 # 最大等待时长,防止无限循环 # 循环加载所有点赞者 while True: try: # 等待"显示更多"按钮可点击 show_more_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "display-flex.p5"))) # 滚动到按钮位置,避免被页面元素遮挡 driver.execute_script("arguments[0].scrollIntoView();", show_more_btn) show_more_btn.click() print("已加载更多点赞者") time.sleep(2) # 给页面加载留缓冲时间 except (NoSuchElementException, ElementClickInterceptedException): # 无更多按钮或无法点击,说明加载完成 print("已加载所有点赞者") break except Exception as e: print(f"加载出错: {str(e)}") time.sleep(3) # 超过最大等待时长则终止 if time.time() - start_time > max_wait_time: print("超过最大等待时间,终止加载") break # 提取已加载的用户基础信息 user_items = likes_container.find_elements(By.CLASS_NAME, "artdeco-list__item") user_info_list = [] for item in user_items: try: name = item.find_element(By.CLASS_NAME, "artdeco-entity-lockup__title").text.strip() profile_link = item.find_element(By.TAG_NAME, "a").get_attribute("href") user_info_list.append({"姓名": name, "主页链接": profile_link}) print(f"已提取用户: {name}") except Exception as e: print(f"提取用户信息出错: {str(e)}") continue # 逐个打开主页获取更多信息(示例:提取职位) for user in user_info_list: try: driver.get(user["主页链接"]) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "text-body-medium"))) position = driver.find_element(By.CLASS_NAME, "text-body-medium").text.strip() user["职位"] = position print(f"{user['姓名']} 的职位: {position}") # 返回点赞弹窗页面,继续处理下一个用户 driver.back() wait.until(EC.presence_of_element_located((By.CLASS_NAME, "artdeco-modal__content"))) time.sleep(1) except Exception as e: print(f"获取{user['姓名']}主页信息出错: {str(e)}") driver.back() continue # 关闭浏览器 driver.quit() # 打印最终结果 print("\n提取的用户信息:") for info in user_info_list: print(info)
关键注意事项
- 反爬规避:LinkedIn对自动化操作检测严格,建议在操作中加入随机等待时间,避免短时间内大量请求触发验证码或账号限制
- 元素更新:LinkedIn的页面元素class名称可能会不定期更新,若代码失效,需重新检查元素的定位符(class/xpath)
- 登录建议:优先手动完成登录后再运行代码,自动登录逻辑容易被平台检测到,导致账号受限
内容的提问来源于stack exchange,提问作者Ashish Kumar
相关产品推荐
相关产品推荐

