Selenium通过XPATH定位元素耗时过长,求经纪人爬取优化方案
如何缩短Selenium爬取经纪人姓名的耗时?
我尝试从bhhs.com的经纪人搜索结果页面爬取经纪人姓名和邮箱,现有代码先抓取首页所有经纪人的个人主页链接,再逐个访问详情页获取姓名和邮箱。但通过XPATH定位包含经纪人姓名的锚标签耗时极长,以下是我的代码:
import os from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.wait import WebDriverWait class MessageIndividual(webdriver.Chrome): def __init__(self, driver_path=r";;C:/SeleniumDriver", teardown=False): self.driver_path = driver_path self.teardown = teardown os.environ['PATH'] += self.driver_path #options = webdriver.ChromeOptions() #options.headless = True super(MessageIndividual, self).__init__() self.implicitly_wait(5) self.maximize_window() def __exit__(self, exc_type, exc_val, exc_tb): if self.teardown: self.quit() def goToSite(self): url = 'https://www.bhhs.com/agent-search-results' self.get(url) def getDetails(self): mylist = [my_elem.get_attribute("href") for my_elem in WebDriverWait(self, 1000).until( EC.visibility_of_all_elements_located((By.XPATH, "//section[@class='cmp-agent-results-list-view']/div[@class='cmp-agent-results-list-view__content container ']/div[@class='row associate pt-3 pb-3 ']/div[@class='col-6 col-sm-4 col-lg-3 order-lg-3 associate__btn-group']/section[2]/a[@href]")))] for i in mylist: self.execute_script("window.open('');") self.switch_to.window(self.window_handles[1]) self.get(i) name = WebDriverWait(self,5).until( EC.presence_of_element_located((By.XPATH,'//h1[@class="cmp-agent__name"]/a[1]')) ) print(name.text) email = WebDriverWait(self,1).until(EC.presence_of_element_located((By.CLASS_NAME,'cmp-agent-details__mail'))) print(email.text) self.close() self.switch_to.window(self.window_handles[0]) if __name__ == '__main__': inst = MessageIndividual(teardown=False) inst.goToSite() inst.getDetails()
优化方案
1. 用CSS选择器替代复杂XPATH
CSS选择器在Selenium中的定位效率普遍比冗长的XPATH更高,把姓名定位的代码改成:
name = WebDriverWait(self,5).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'h1.cmp-agent__name > a:first-child')) )
首页的链接定位也可以简化为CSS选择器,减少路径复杂度:
mylist = [my_elem.get_attribute("href") for my_elem in WebDriverWait(self, 10).until( EC.visibility_of_all_elements_located((By.CSS_SELECTOR, 'section.cmp-agent-results-list-view .associate__btn-group section:nth-child(2) a[href]')) )]
2. 移除隐式等待,只保留显式等待
代码里的implicitly_wait(5)会和显式等待冲突,导致所有元素查找都额外等待,直接删掉这行,只在需要的元素上用显式等待控制超时时间。
3. 开启Headless模式
取消代码中Headless相关注释,启用无界面模式,跳过页面渲染能大幅提速:
options = webdriver.ChromeOptions() options.add_argument('--headless=new') options.add_argument('--disable-gpu') super(MessageIndividual, self).__init__(options=options)
4. 直接从首页提取姓名(最有效的优化)
观察首页结构,经纪人姓名已经显示在列表里,完全不用进详情页拿,这能省掉大量页面加载时间。修改首页数据提取逻辑:
def getDetails(self): # 同时获取姓名和链接 agent_items = WebDriverWait(self, 10).until( EC.visibility_of_all_elements_located((By.CSS_SELECTOR, '.associate__info h3 a')) ) agent_links = [item.get_attribute("href") for item in agent_items] agent_names = [item.text for item in agent_items] # 再遍历链接拿邮箱 for link, name in zip(agent_links, agent_names): print(name) self.execute_script("window.open('');") self.switch_to.window(self.window_handles[1]) self.get(link) email = WebDriverWait(self,1).until(EC.presence_of_element_located((By.CLASS_NAME,'cmp-agent-details__mail'))) print(email.text) self.close() self.switch_to.window(self.window_handles[0])
5. 优化窗口切换逻辑
每次打开新窗口再切换的开销不小,可以改成直接在当前页面跳转后返回:
for link, name in zip(agent_links, agent_names): print(name) self.get(link) email = WebDriverWait(self,1).until(EC.presence_of_element_located((By.CLASS_NAME,'cmp-agent-details__mail'))) print(email.text) self.back()
不过这种方式要注意页面缓存问题,稳定性可能不如新窗口,但能减少切换开销。
内容的提问来源于stack exchange,提问作者Junaid Nazir
相关产品推荐
相关产品推荐

