You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium脚本搜索https://ssllc.com/不稳定问题求助

修复Selenium搜索脚本的可靠性问题

核心问题分析

  1. 搜索词格式错误:用+替代空格作为分隔符,网站搜索框实际需要空格分隔关键词,导致长搜索词无法匹配结果。
  2. Headless模式被检测:默认headless配置易被识别为爬虫,引发加载异常或无结果返回。
  3. 等待逻辑不完善:仅等待结果容器存在,未确保结果项加载完成;固定超时时间无法适配网站负载波动。
  4. CSS选择器不稳定:依赖Gatsby生成的动态ID(如#gatsby-focus-wrapper),页面结构变化会直接导致选择器失效。
  5. 浏览器实例重复初始化:每次搜索重启浏览器,既降低效率又容易触发反爬机制。

针对性修复方案

  • 修正搜索词格式:将搜索词中的+替换为空格,匹配网站原生搜索逻辑。
  • 优化Headless配置:添加模拟正常浏览器的参数,规避反爬检测。
  • 改进等待逻辑:等待结果项加载完成,延长超时时间并增加重试机制应对高负载场景。
  • 使用稳定选择器:基于页面固定类名(如.ais-Hits)定位元素,摆脱动态ID依赖。
  • 复用浏览器实例:单个浏览器实例完成所有搜索,提升效率并降低反爬风险。

修改后的完整代码

from selenium.webdriver import Chrome
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options as ChromeOptions
from selenium.common.exceptions import TimeoutException, NoSuchElementException

# 基础配置
base_url = 'https://www.ssllc.com'
# 更稳定的搜索框选择器(基于属性+类名)
search_bar_selector = 'input[type="text"][placeholder="Search..."]'
# 直接定位结果项,而非整个列表容器
result_item_selector = '.ais-Hits ul li'

# 修正后的搜索词(替换+为空格)
search_queries = [
    'Unused Sartorius 1000 Liter BIOSTAT CultiBag STR Single Use Bioreactor',
    '3 x V5/XCell Repigen Next Gen ATF controllers',
    'InSite Integrity Tester'
]

# 优化Chrome配置,模拟正常浏览器
options = ChromeOptions()
options.add_argument('--headless=new')  # 新版headless模式更难被检测
options.add_argument('--disable-blink-features=AutomationControlled')
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
options.add_argument('--window-size=1920,1080')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')

def perform_search(driver, query):
    max_retries = 3
    for attempt in range(max_retries):
        try:
            # 回到首页清除残留搜索状态
            driver.get(base_url)
            
            # 等待搜索框并输入查询
            search_input = WebDriverWait(driver, 15).until(
                EC.visibility_of_element_located((By.CSS_SELECTOR, search_bar_selector))
            )
            search_input.clear()
            search_input.send_keys(query)
            search_input.send_keys(Keys.RETURN)
            
            # 等待结果项加载完成,延长超时时间适配高负载
            WebDriverWait(driver, 20).until(
                EC.presence_of_element_located((By.CSS_SELECTOR, result_item_selector))
            )
            
            # 提取所有结果详情
            results = driver.find_elements(By.CSS_SELECTOR, result_item_selector)
            if results:
                print(f"【{query}】搜索结果:")
                for idx, result in enumerate(results, 1):
                    title = result.find_element(By.CSS_SELECTOR, 'h3').text.strip()
                    link = result.find_element(By.TAG_NAME, 'a').get_attribute('href')
                    print(f"{idx}. {title}\n链接:{link}\n")
            else:
                print(f"【{query}】未找到相关结果\n")
            break
        except (TimeoutException, NoSuchElementException) as e:
            if attempt == max_retries - 1:
                print(f"【{query}】搜索失败(已重试{max_retries}次):{str(e)}\n")
            else:
                print(f"【{query}】第{attempt+1}次搜索超时,正在重试...")

try:
    # 复用浏览器实例完成所有搜索
    with Chrome(options=options) as driver:
        for query in search_queries:
            perform_search(driver, query)
except Exception as e:
    print(f"全局错误:{str(e)}")

额外说明

  • 新增的重试机制可应对网站高负载导致的超时问题,最多重试3次。
  • 替换后的选择器基于页面静态类名和属性,不受Gatsby动态ID变化影响。
  • 模拟真实用户的UA和窗口尺寸,降低被反爬机制拦截的概率。

内容的提问来源于stack exchange,提问作者Aaron S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 04:15:19