Selenium脚本搜索https://ssllc.com/不稳定问题求助
修复Selenium搜索脚本的可靠性问题
核心问题分析
- 搜索词格式错误:用
+替代空格作为分隔符,网站搜索框实际需要空格分隔关键词,导致长搜索词无法匹配结果。 - Headless模式被检测:默认headless配置易被识别为爬虫,引发加载异常或无结果返回。
- 等待逻辑不完善:仅等待结果容器存在,未确保结果项加载完成;固定超时时间无法适配网站负载波动。
- CSS选择器不稳定:依赖Gatsby生成的动态ID(如
#gatsby-focus-wrapper),页面结构变化会直接导致选择器失效。 - 浏览器实例重复初始化:每次搜索重启浏览器,既降低效率又容易触发反爬机制。
针对性修复方案
- 修正搜索词格式:将搜索词中的
+替换为空格,匹配网站原生搜索逻辑。 - 优化Headless配置:添加模拟正常浏览器的参数,规避反爬检测。
- 改进等待逻辑:等待结果项加载完成,延长超时时间并增加重试机制应对高负载场景。
- 使用稳定选择器:基于页面固定类名(如
.ais-Hits)定位元素,摆脱动态ID依赖。 - 复用浏览器实例:单个浏览器实例完成所有搜索,提升效率并降低反爬风险。
修改后的完整代码
from selenium.webdriver import Chrome from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options as ChromeOptions from selenium.common.exceptions import TimeoutException, NoSuchElementException # 基础配置 base_url = 'https://www.ssllc.com' # 更稳定的搜索框选择器(基于属性+类名) search_bar_selector = 'input[type="text"][placeholder="Search..."]' # 直接定位结果项,而非整个列表容器 result_item_selector = '.ais-Hits ul li' # 修正后的搜索词(替换+为空格) search_queries = [ 'Unused Sartorius 1000 Liter BIOSTAT CultiBag STR Single Use Bioreactor', '3 x V5/XCell Repigen Next Gen ATF controllers', 'InSite Integrity Tester' ] # 优化Chrome配置,模拟正常浏览器 options = ChromeOptions() options.add_argument('--headless=new') # 新版headless模式更难被检测 options.add_argument('--disable-blink-features=AutomationControlled') options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') options.add_argument('--window-size=1920,1080') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') def perform_search(driver, query): max_retries = 3 for attempt in range(max_retries): try: # 回到首页清除残留搜索状态 driver.get(base_url) # 等待搜索框并输入查询 search_input = WebDriverWait(driver, 15).until( EC.visibility_of_element_located((By.CSS_SELECTOR, search_bar_selector)) ) search_input.clear() search_input.send_keys(query) search_input.send_keys(Keys.RETURN) # 等待结果项加载完成,延长超时时间适配高负载 WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.CSS_SELECTOR, result_item_selector)) ) # 提取所有结果详情 results = driver.find_elements(By.CSS_SELECTOR, result_item_selector) if results: print(f"【{query}】搜索结果:") for idx, result in enumerate(results, 1): title = result.find_element(By.CSS_SELECTOR, 'h3').text.strip() link = result.find_element(By.TAG_NAME, 'a').get_attribute('href') print(f"{idx}. {title}\n链接:{link}\n") else: print(f"【{query}】未找到相关结果\n") break except (TimeoutException, NoSuchElementException) as e: if attempt == max_retries - 1: print(f"【{query}】搜索失败(已重试{max_retries}次):{str(e)}\n") else: print(f"【{query}】第{attempt+1}次搜索超时,正在重试...") try: # 复用浏览器实例完成所有搜索 with Chrome(options=options) as driver: for query in search_queries: perform_search(driver, query) except Exception as e: print(f"全局错误:{str(e)}")
额外说明
- 新增的重试机制可应对网站高负载导致的超时问题,最多重试3次。
- 替换后的选择器基于页面静态类名和属性,不受Gatsby动态ID变化影响。
- 模拟真实用户的UA和窗口尺寸,降低被反爬机制拦截的概率。
内容的提问来源于stack exchange,提问作者Aaron S
相关产品推荐
相关产品推荐

