You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取LinkedIn公司名时循环重复输出首个名称问题

问题原因与解决办法

问题根源

你循环中使用的//h4是绝对路径XPath,它会从整个页面的根节点开始搜索匹配元素,所以每次循环都会返回页面中第一个<h4>标签的内容,导致重复输出第一个公司名称。另外,直接通过By.TAG_NAME, 'li'获取的列表包含大量页面无关的<li>元素,进一步干扰结果。

修复方案

  1. 使用相对路径XPath:在遍历单个职位元素时,用.//h4(或更精确的带class定位)限定搜索范围为当前职位元素的子节点。
  2. 精准定位职位卡片:直接定位LinkedIn职位卡片的<li>元素(class为base-search-card),避免遍历无关元素。
  3. 添加显式等待:确保页面元素加载完成后再进行抓取,避免动态加载导致的元素缺失。

修改后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

url = 'https://www.linkedin.com/jobs/search/?currentJobId=3492578215&geoId=105365761&keywords=data%20analyst&location=Nigeria&refresh=true'
options = webdriver.ChromeOptions()
options.add_experimental_option('detach', True)
driver = webdriver.Chrome(r"C:\Users\i\Desktop\PPstuff\selenium\chromedriver.exe", options=options)
driver.get(url)

# 等待职位卡片加载完成
wait = WebDriverWait(driver, 10)
jobs = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'base-search-card')))

company_name = []
for job in jobs:
    # 使用相对路径定位当前职位下的公司名称
    company = job.find_element(By.XPATH, ".//h4[@class='base-search-card__subtitle']").text
    company_name.append(company)
    print(company)

# 转换为DataFrame
df = pd.DataFrame({'公司名称': company_name})
print(df)

额外提示

  • LinkedIn有反爬机制,频繁抓取可能触发限制,建议添加随机延时(time.sleep(random.uniform(1,3)))。
  • 若需要抓取更多职位,需处理页面滚动加载的逻辑,模拟滚动到底部加载新内容。

内容的提问来源于stack exchange,提问作者Dayo Salam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 03:24:58