You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium的LinkedIn爬虫速度优化求助

目标
  • 我正在开发一个基于Python的LinkedIn网络爬虫,程序接收用户登录凭证和自定义搜索查询作为输入,通过Selenium导航到搜索结果中的个人资料页面,提取指定网页元素的数据并存储到Pandas DataFrame中。
问题
  • 以下是我的代码,目前程序运行耗时过长(解析68个个人资料约需23分钟),恳请帮助优化代码提升运行速度,谢谢!
代码
# imports
import time
from selenium import webdriver
from selenium.webdriver.common.by import By
import pandas as pd

userid = "userid@domain.com"
password = "p@Ssw0rd!"
keyword = "Master of Business Data Science Otago "
url = f"https://www.linkedin.com/search/results/people/?keywords={keyword}&origin=SWITCH_SEARCH_VERTICAL&sid=RZW"

driver = webdriver.Chrome()

driver.get("https://www.linkedin.com")
driver.implicitly_wait(6)
driver.find_element(By.XPATH, """//*[@id="session_key"]""").send_keys(userid)
driver.find_element(By.XPATH, """//*[@id="session_password"]""").send_keys(password)
driver.find_element(By.XPATH, "//button[@class='sign-in-form__submit-button']").click()

driver.get(url)
links = []
scroll_target = driver.find_element(By.CLASS_NAME,"background-mercado") # target for scrolling page to linked in logo at the bottom so that the 'Next' button's element becomes visible
driver.execute_script('arguments[0].scrollIntoView(true)',scroll_target)

while True:
    try:
        time.sleep(3) 
        linky = driver.find_elements(By.CLASS_NAME,'app-aware-link ') # locating containers housing links in the page
        links.append([li.get_attribute('href') for li in linky[::2] if 'miniProfileUrn'in str(li.get_attribute('href'))]) # filtering only profile links from list of all links
        page_button = driver.find_element(By.XPATH, '//button[@aria-label="Next"]') # locating the 'Next' button to click to the next page of results
        page_button.click()
    except:
        print("No more pages")

# locating and parsing individual elements in each profile for profile links captured
elements = { 
    'name': """//h1[@class="text-heading-xlarge inline t-24 v-align-middle break-words"]""",
    'prefix':"""//div[@class="text-body-small v-align-middle break-words t-black--light"]""",
    'title':"""//div[@class="text-body-medium break-words"]""",
    'location':"""//div[@class="text-body-small inline t-black--light break-words"]""",
    'see_more':"""//button[@class="inline-show-more-text__button
                inline-show-more-text__button--light
                link"]""",
    'about':"""//div[@class="inline-show-more-text




         full-width"]""",
    'experience':"""//ul[@class="pvs-list


            "]""",
    'expander':"""a[@class="optional-action-target-wrapper artdeco-button artdeco-button--tertiary artdeco-button--standard artdeco-button--2 artdeco-button--muted 
          inline-flex justify-center full-width align-items-center artdeco-button--fluid
          
          "]"""
}
profiles = []
for n in links:
    for m in n:
        driver.get(m)
        time.sleep(2)
        try:
            see_mores = driver.find_elements(By.XPATH, elements['see_more']) # locating 'see more' button at paragraph ends for long descriptive fields to expand them
            for s in see_mores:
                s.click()
        except:
            print("No 'see more' button")
        time.sleep(1)
        try:
            expanders = driver.find_elements(By.XPATH, elements['expander']) # locating expansion buttons at to expand sections with collapsed data entries
            for e in expanders:
                e.click()
        except:
            print("No expanders")
        time.sleep(1)
        try:
            name = driver.find_element(By.XPATH,elements['name']).text
        except:
            name = None
        try:
            prefix = driver.find_element(By.XPATH,elements['prefix']).text
        except:
            prefix = None
        try:
            title = driver.find_element(By.XPATH,elements['title']).text
        except:
            title= None
        try:
            location = driver.find_element(By.XPATH,elements['location']).text
        except:
            location = None
        try:
            about = driver.find_element(By.XPATH,elements['about']).text
        except:
            about = None
        try:    
            experience = driver.find_elements(By.XPATH,elements['experience'])[1].text
        except:
            experience = None
        profiles.append({'name':name,'prefix':prefix,'title':title,'about':about,'experience':experience})

# storing to a data frame and exporting to a csv
profile_df = pd.DataFrame(profiles)
profile_df.to_csv("D:\\linkedin_profiles.csv")
优化方案

1. 替换固定休眠为显式等待

固定time.sleep()是最大的耗时来源,改用WebDriverWait等待元素加载完成再操作,避免无意义的等待:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 示例:等待"Next"按钮可点击后再点击
page_button = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.XPATH, '//button[@aria-label="Next"]'))
)
page_button.click()

# 示例:等待个人资料页面的名称元素加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, elements['name']))
)

2. 启用无头浏览器模式

关闭Chrome的UI渲染,减少资源开销:

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--disable-gpu")
options.add_argument("--window-size=1920,1080")
driver = webdriver.Chrome(options=options)

3. 优化元素定位逻辑

  • 优先使用By.CSS_SELECTOR替代冗长的XPATH,定位效率更高,例如名称元素可简化为By.CSS_SELECTOR, "h1.text-heading-xlarge"
  • 修复elements字典中XPATH的格式问题(多行空格会导致定位失效),确保表达式简洁准确

4. 并行处理个人资料

使用多线程同时处理多个个人资料页面(注意控制并发数,避免触发LinkedIn反爬):

from concurrent.futures import ThreadPoolExecutor

def scrape_profile(profile_url):
    temp_options = webdriver.ChromeOptions()
    temp_options.add_argument("--headless=new")
    temp_driver = webdriver.Chrome(options=temp_options)
    temp_driver.get(profile_url)
    
    # 执行元素提取逻辑(复用原有的提取代码)
    # ...
    
    temp_driver.quit()
    return profile_data

# 先把嵌套的links列表扁平化
flat_links = [link for sublist in links for link in sublist]
# 控制并发数为3,避免账号受限
with ThreadPoolExecutor(max_workers=3) as executor:
    results = executor.map(scrape_profile, flat_links)
profiles = list(results)

5. 优化链接收集流程

  • 滚动加载当前页所有结果后再一次性提取链接,减少页面切换等待:
# 滚动加载当前页全部内容
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(1)
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height
# 提取当前页所有有效链接
linky = driver.find_elements(By.CLASS_NAME,'app-aware-link ')
current_links = [li.get_attribute('href') for li in linky if 'miniProfileUrn' in str(li.get_attribute('href'))]
links.extend(current_links)
  • 收集链接时直接去重,避免重复处理同一 profile

6. 减少页面操作冗余

  • 点击展开按钮前,先检查元素是否存在且可点击,避免无效的异常捕获和等待
  • 一次性获取所有需要的元素,减少多次DOM查询的开销

7. 额外优化项

  • 禁用图片加载,减少网络请求:
options.add_argument("--blink-settings=imagesEnabled=false")
  • 使用新标签页打开个人资料,处理完成后关闭标签页,减少页面重载时间

内容的提问来源于stack exchange,提问作者Richard T Vetticad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 21:39:36