You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取LinkedIn经历板块中就职公司名称问题求助

解决方案

核心问题分析

  1. 原代码错误地将外层元素定位为div,但实际目标元素的外层是span,导致无法匹配到正确节点。
  2. 使用find_element(单数方法)只会获取第一个匹配元素,无法遍历所有公司条目。
  3. 未处理动态加载:LinkedIn页面内容是动态渲染的,直接获取元素可能出现未加载完成的情况。

修正后的代码示例

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待页面加载完成,获取所有包含公司名称的span元素
wait = WebDriverWait(driver, 10)
# 定位所有符合结构的目标span(外层span的class是t-14 t-normal,内部包含aria-hidden="true"的span)
company_span_list = wait.until(
    EC.presence_of_all_elements_located(
        (By.CSS_SELECTOR, 'span.t-14.t-normal span[aria-hidden="true"]')
    )
)

# 遍历提取并清理公司名称
company_names = []
for span in company_span_list:
    raw_content = span.text.strip()
    # 去除文本首尾的引号
    if raw_content.startswith('"') and raw_content.endswith('"'):
        raw_content = raw_content[1:-1]
    # 分割出公司名称(按第一个句号分割,去掉后面的任职类型等内容)
    company_name = raw_content.split('.')[0].strip()
    if company_name:
        company_names.append(company_name)

# 输出所有公司名称
for name in company_names:
    print(name)

关键注意事项

  • 登录验证:LinkedIn需要登录才能查看用户完整经历,确保你的WebDriver已经完成登录流程(可以手动登录或通过代码注入cookie)。
  • 反爬规避:不要频繁刷新或请求,建议在遍历操作中加入短时间等待(如time.sleep(0.5)),避免触发反爬机制导致账号受限。
  • 元素结构变化:LinkedIn的页面结构可能随时更新,如果后续出现定位失败,需要重新检查目标元素的CSS选择器或XPath。

内容的提问来源于stack exchange,提问作者maven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 01:46:30