Docker部署的FastAPI+Selenium汽车品牌价格查询项目无法获取Autovit.ro的品牌与价格元素
Docker部署的FastAPI+Selenium汽车品牌价格查询项目无法获取Autovit.ro的品牌与价格元素
嘿,我完全懂你现在的困扰——Autovit.ro的页面元素类名大多是动态生成的(就是那些ooa-xxx、e4zbkti0这类看起来随机的类),每次页面更新甚至不同会话都会变,用这些类名抓内容肯定会扑空。咱们一步步来解决这个问题:
核心问题:替换不稳定的元素选择器
首先要抛弃依赖动态类名的写法,改用基于元素结构、属性或者语义的稳定选择器:
1. 修复品牌列表抓取逻辑
Autovit的品牌列表在顶部的「Marca」下拉框里,咱们可以通过等待下拉框加载并展开,再抓取里面的选项:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def get_car_brands(): url = "https://www.autovit.ro/" driver = get_selenium_driver() driver.get(url) brands = [] try: # 等待品牌下拉容器加载并点击展开 brand_dropdown = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//div[contains(text(), 'Marca')]/parent::div")) ) brand_dropdown.click() # 等待品牌列表完全加载 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.XPATH, "//ul[contains(@class, 'search-dropdown__list')]/li/a")) ) soup = BeautifulSoup(driver.page_source, 'html.parser') # 提取品牌文本,排除"Toate marcarile"选项 for brand in soup.select("ul.search-dropdown__list li a"): brand_text = brand.text.strip() if brand_text and brand_text != "Toate marcarile": brands.append(brand_text) except Exception as e: print(f"提取品牌出错: {e}") driver.quit() return brands
2. 修复价格范围抓取逻辑
汽车价格通常带有data-testid="ad-price"属性,咱们可以直接用这个属性定位,同时过滤掉「Pret negociabil」这类非明确价格的内容:
def get_price_range(brand): # 构造正确的品牌URL:把品牌名转成小写+连字符格式 brand_slug = brand.lower().replace(" ", "-") url = f"https://www.autovit.ro/marca/{brand_slug}" driver = get_selenium_driver() driver.get(url) prices = [] try: # 等待价格元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.XPATH, "//span[contains(@data-testid, 'ad-price')]")) ) soup = BeautifulSoup(driver.page_source, 'html.parser') # 提取并清洗价格 for price_tag in soup.select("span[data-testid='ad-price']"): price_text = price_tag.text.strip() # 跳过可议价的条目,只处理带€的明确价格 if "€" in price_text and "negociabil" not in price_text.lower(): # 清理价格字符串:去掉€、空格、标点 clean_price = price_text.replace("€", "").replace(" ", "").replace(".", "").replace(",", "") try: prices.append(int(clean_price)) except ValueError: pass # 跳过无法转换的异常价格 except Exception as e: print(f"提取价格出错: {e}") driver.quit() return min(prices), max(prices) if prices else (0, 0)
额外优化:解决反爬与Docker环境问题
1. 优化Chrome配置,避免被反爬检测
Autovit会检测headless浏览器,咱们给Chrome加一些参数绕过检测:
def get_selenium_driver(): options = webdriver.ChromeOptions() options.headless = True # 解决Docker环境下的沙箱问题 options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") # 绕过自动化检测 options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) # 设置窗口大小,避免布局错乱 options.add_argument("--window-size=1920,1080") # 如果是Docker环境,直接用系统安装的chromedriver(路径参考Dockerfile) driver = webdriver.Chrome(service=Service("/usr/bin/chromedriver"), options=options) # 执行JS隐藏webdriver标识 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") return driver
2. Docker环境配置要点
在Docker里运行Chrome需要安装依赖,给你一个示例Dockerfile:
FROM python:3.10-slim WORKDIR /app # 复制依赖文件并安装 COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt # 安装Chrome、chromedriver及系统依赖 RUN apt-get update && apt-get install -y \ chromium \ chromium-driver \ libglib2.0-0 \ libnss3 \ libfontconfig1 \ && rm -rf /var/lib/apt/lists/* # 将系统chromedriver加入环境变量 ENV PATH="/usr/bin:${PATH}" # 复制项目代码 COPY . . # 启动FastAPI服务 CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
对应的requirements.txt内容:
fastapi uvicorn selenium beautifulsoup4
最后提醒
- 尽量用
WebDriverWait替代time.sleep,等待元素加载完成再操作,比固定延时更可靠 - Autovit的页面结构可能会更新,如果之后又抓不到内容,记得重新检查元素的定位方式
- 不要太频繁请求,避免被网站封禁IP
备注:内容来源于stack exchange,提问作者Dobrea Marian
相关产品推荐
相关产品推荐

