基于Selenium的LinkedIn爬虫速度优化求助
目标
- 我正在开发一个基于Python的LinkedIn网络爬虫,程序接收用户登录凭证和自定义搜索查询作为输入,通过Selenium导航到搜索结果中的个人资料页面,提取指定网页元素的数据并存储到Pandas DataFrame中。
问题
- 以下是我的代码,目前程序运行耗时过长(解析68个个人资料约需23分钟),恳请帮助优化代码提升运行速度,谢谢!
代码
# imports import time from selenium import webdriver from selenium.webdriver.common.by import By import pandas as pd userid = "userid@domain.com" password = "p@Ssw0rd!" keyword = "Master of Business Data Science Otago " url = f"https://www.linkedin.com/search/results/people/?keywords={keyword}&origin=SWITCH_SEARCH_VERTICAL&sid=RZW" driver = webdriver.Chrome() driver.get("https://www.linkedin.com") driver.implicitly_wait(6) driver.find_element(By.XPATH, """//*[@id="session_key"]""").send_keys(userid) driver.find_element(By.XPATH, """//*[@id="session_password"]""").send_keys(password) driver.find_element(By.XPATH, "//button[@class='sign-in-form__submit-button']").click() driver.get(url) links = [] scroll_target = driver.find_element(By.CLASS_NAME,"background-mercado") # target for scrolling page to linked in logo at the bottom so that the 'Next' button's element becomes visible driver.execute_script('arguments[0].scrollIntoView(true)',scroll_target) while True: try: time.sleep(3) linky = driver.find_elements(By.CLASS_NAME,'app-aware-link ') # locating containers housing links in the page links.append([li.get_attribute('href') for li in linky[::2] if 'miniProfileUrn'in str(li.get_attribute('href'))]) # filtering only profile links from list of all links page_button = driver.find_element(By.XPATH, '//button[@aria-label="Next"]') # locating the 'Next' button to click to the next page of results page_button.click() except: print("No more pages") # locating and parsing individual elements in each profile for profile links captured elements = { 'name': """//h1[@class="text-heading-xlarge inline t-24 v-align-middle break-words"]""", 'prefix':"""//div[@class="text-body-small v-align-middle break-words t-black--light"]""", 'title':"""//div[@class="text-body-medium break-words"]""", 'location':"""//div[@class="text-body-small inline t-black--light break-words"]""", 'see_more':"""//button[@class="inline-show-more-text__button inline-show-more-text__button--light link"]""", 'about':"""//div[@class="inline-show-more-text full-width"]""", 'experience':"""//ul[@class="pvs-list "]""", 'expander':"""a[@class="optional-action-target-wrapper artdeco-button artdeco-button--tertiary artdeco-button--standard artdeco-button--2 artdeco-button--muted inline-flex justify-center full-width align-items-center artdeco-button--fluid "]""" } profiles = [] for n in links: for m in n: driver.get(m) time.sleep(2) try: see_mores = driver.find_elements(By.XPATH, elements['see_more']) # locating 'see more' button at paragraph ends for long descriptive fields to expand them for s in see_mores: s.click() except: print("No 'see more' button") time.sleep(1) try: expanders = driver.find_elements(By.XPATH, elements['expander']) # locating expansion buttons at to expand sections with collapsed data entries for e in expanders: e.click() except: print("No expanders") time.sleep(1) try: name = driver.find_element(By.XPATH,elements['name']).text except: name = None try: prefix = driver.find_element(By.XPATH,elements['prefix']).text except: prefix = None try: title = driver.find_element(By.XPATH,elements['title']).text except: title= None try: location = driver.find_element(By.XPATH,elements['location']).text except: location = None try: about = driver.find_element(By.XPATH,elements['about']).text except: about = None try: experience = driver.find_elements(By.XPATH,elements['experience'])[1].text except: experience = None profiles.append({'name':name,'prefix':prefix,'title':title,'about':about,'experience':experience}) # storing to a data frame and exporting to a csv profile_df = pd.DataFrame(profiles) profile_df.to_csv("D:\\linkedin_profiles.csv")
优化方案
1. 替换固定休眠为显式等待
固定time.sleep()是最大的耗时来源,改用WebDriverWait等待元素加载完成再操作,避免无意义的等待:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 示例:等待"Next"按钮可点击后再点击 page_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[@aria-label="Next"]')) ) page_button.click() # 示例:等待个人资料页面的名称元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, elements['name'])) )
2. 启用无头浏览器模式
关闭Chrome的UI渲染,减少资源开销:
options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument("--disable-gpu") options.add_argument("--window-size=1920,1080") driver = webdriver.Chrome(options=options)
3. 优化元素定位逻辑
- 优先使用
By.CSS_SELECTOR替代冗长的XPATH,定位效率更高,例如名称元素可简化为By.CSS_SELECTOR, "h1.text-heading-xlarge" - 修复
elements字典中XPATH的格式问题(多行空格会导致定位失效),确保表达式简洁准确
4. 并行处理个人资料
使用多线程同时处理多个个人资料页面(注意控制并发数,避免触发LinkedIn反爬):
from concurrent.futures import ThreadPoolExecutor def scrape_profile(profile_url): temp_options = webdriver.ChromeOptions() temp_options.add_argument("--headless=new") temp_driver = webdriver.Chrome(options=temp_options) temp_driver.get(profile_url) # 执行元素提取逻辑(复用原有的提取代码) # ... temp_driver.quit() return profile_data # 先把嵌套的links列表扁平化 flat_links = [link for sublist in links for link in sublist] # 控制并发数为3,避免账号受限 with ThreadPoolExecutor(max_workers=3) as executor: results = executor.map(scrape_profile, flat_links) profiles = list(results)
5. 优化链接收集流程
- 滚动加载当前页所有结果后再一次性提取链接,减少页面切换等待:
# 滚动加载当前页全部内容 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(1) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 提取当前页所有有效链接 linky = driver.find_elements(By.CLASS_NAME,'app-aware-link ') current_links = [li.get_attribute('href') for li in linky if 'miniProfileUrn' in str(li.get_attribute('href'))] links.extend(current_links)
- 收集链接时直接去重,避免重复处理同一 profile
6. 减少页面操作冗余
- 点击展开按钮前,先检查元素是否存在且可点击,避免无效的异常捕获和等待
- 一次性获取所有需要的元素,减少多次DOM查询的开销
7. 额外优化项
- 禁用图片加载,减少网络请求:
options.add_argument("--blink-settings=imagesEnabled=false")
- 使用新标签页打开个人资料,处理完成后关闭标签页,减少页面重载时间
内容的提问来源于stack exchange,提问作者Richard T Vetticad
相关产品推荐
相关产品推荐

