You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium抓取Thriveil网站产品信息时元素检测失败的求助

使用Selenium抓取Thriveil网站产品信息时元素检测失败的求助

Hey there! I took a close look at your code and the problem you're facing with element detection on the Thriveil site. Let's break down the key issues and fix them step by step—no jargon, just practical changes to get your scraper working reliably.


1. Critical Bug: Incorrect Data Collection Logic

Right now, your code is appending entire lists (like names, brand_names) to each product entry, which is not only wrong but can also cause unexpected behavior. Instead, you should collect individual product data for each item and add that single entry to your product_list.

For example, instead of:

product_data = {
    "name": names,
    "brand_name": brand_names,
    ...
}
product_list.append(product_data)

You should do:

# For each product, create a single dict with its data
product_data = {
    "name": name.text.strip(),
    "brand_name": brand.text.strip(),
    "brand_link": brand_element.get_attribute('href'),
    "strain": strain.text.strip() if strain else "N/A",
    "potency": potency_text if 'potency_text' in locals() else "N/A",
    "price": price.text.strip() if price else "N/A",
    "effects": ", ".join([e.text.strip() for e in effect_elements]) if effect_elements else "N/A"
}
product_list.append(product_data)

And you can remove the global lists (names, brand_names, etc.) entirely—they're not needed anymore.


2. Use Targeted Waits for Every Element

Waiting just for the <h1> isn't enough. Many elements load asynchronously after the page initializes. For each element you want to scrape, use WebDriverWait to ensure it's present before trying to access it. This will fix most "element not found" errors.

Example for scraping the product name:

# Wait for the product name to be present
name = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "h1[data-testid='product-name']"))).text.strip()

Do this for every element (brand, strain, price, etc.) instead of using driver.find_element directly.


3. Avoid Fragile Auto-Generated Class Names

Class names like full-card_Wrapper-sc-11z5u35-0 or typography__Brand-sc-1q7gvs8-2.fyoohd are auto-generated by frameworks like React/Vue and can change anytime the site updates. Instead, use more stable selectors:

  • Data attributes: The site uses data-testid (like data-testid='product-name') which is designed for testing/scraping and rarely changes.
  • Nested selectors: If data attributes aren't available, use parent-child relationships (e.g., div.product-brand a instead of a random class).

For the brand element, try this instead of the fragile class:

brand_element = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='Brand'] a")))
brand_name = brand_element.text.strip()
brand_link = brand_element.get_attribute('href')

4. Bypass Selenium Detection

Many modern sites block Selenium by detecting its unique browser properties. To fix this, add stealth settings to your Edge options to make your scraper look like a regular user:

edge_options = Options()
# Add anti-detection flags
edge_options.add_argument("--disable-blink-features=AutomationControlled")
edge_options.add_experimental_option("excludeSwitches", ["enable-automation"])
edge_options.add_experimental_option('useAutomationExtension', False)
# Set a real user-agent (you can get yours from your browser's dev tools)
edge_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36 Edg/120.0.0.0")

This will help the site not recognize you're using Selenium.


5. Fix Potency Extraction (Avoid Index Errors)

Your code uses potency_values[1] which will crash if there are fewer than 2 potency elements. Add a check to handle this:

potency_text = "N/A"
potency_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.info-chip__InfoChipText-sc-11n9ujc-0")))
if len(potency_elements) >=2:
    potency_text = potency_elements[1].text.split(":")[-1].strip()

Revised Code Snippet (Key Fixes Included)

Here's a trimmed version of your code with all the above fixes applied:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.edge.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.edge.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, TimeoutException
import pandas as pd
import time

# Set up Edge with anti-detection
edge_options = Options()
edge_options.add_argument("--disable-blink-features=AutomationControlled")
edge_options.add_experimental_option("excludeSwitches", ["enable-automation"])
edge_options.add_experimental_option('useAutomationExtension', False)
edge_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36 Edg/120.0.0.0")

service = Service("C:\\Users\\iyush\\Documents\\VS Code\\Selenium\\msedgedriver.exe")  
driver = webdriver.Edge(service=service, options=edge_options)  
url = "https://thriveil.com/casey-rec-menu/?dtche%5Bpath%5D=products"
driver.get(url)  
wait = WebDriverWait(driver, 30)

# Close ad and cookies
try:
    close_button = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "terpli-close")))
    close_button.click()
except TimeoutException:
    print("Ad close button not found. Continuing...")

try:
    accept_button = wait.until(EC.element_to_be_clickable((By.ID, "wt-cli-accept-all-btn")))
    accept_button.click()
except TimeoutException:
    print("Cookie consent button not found. Continuing...")

product_list = []
current_page = 1

while True:
    # Wait for product cards to load
    wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div[class*='full-card_Wrapper']")))
    product_cards = driver.find_elements(By.CSS_SELECTOR, "div[class*='full-card_Wrapper'] a")
    product_urls = [card.get_attribute("href") for card in product_cards if card.get_attribute("href")]

    for url in product_urls:
        driver.execute_script(f"window.open('{url}', '_blank');")
        driver.switch_to.window(driver.window_handles[-1])
        
        try:
            # Wait for critical elements to load
            wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "h1[data-testid='product-name']")))
            
            # Scrape each field with waits
            name = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "h1[data-testid='product-name']"))).text.strip()
            
            brand_element = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='Brand'] a")))
            brand_name = brand_element.text.strip()
            brand_link = brand_element.get_attribute("href")
            
            strain = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "span[data-testid='info-chip']"))).text.strip()
            
            # Potency handling
            potency_text = "N/A"
            potency_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.info-chip__InfoChipText-sc-11n9ujc-0")))
            if len(potency_elements) >=2:
                potency_text = potency_elements[1].text.split(":")[-1].strip()
            
            price = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='price__PriceText']"))).text.strip()
            
            effects = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.effect-tile__Text-sc-1as4rkm-1")))
            effect_text = ", ".join([e.text.strip() for e in effects]) if effects else "N/A"
            
            # Add single product data to list
            product_list.append({
                "name": name,
                "brand_name": brand_name,
                "brand_link": brand_link,
                "strain": strain,
                "potency": potency_text,
                "price": price,
                "effects": effect_text
            })
            
        except Exception as e:
            print(f"Error scraping {url}: {str(e)}")
        finally:
            driver.close()
            driver.switch_to.window(driver.window_handles[0])
    
    print(f"Page {current_page} scraped successfully.")
    
    # Next page handling
    try:
        next_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button[aria-label*='next page']")))
        # Check if next button is disabled
        if "disabled" in next_btn.get_attribute("class"):
            break
        next_btn.click()
        current_page +=1
        time.sleep(2) # Short wait for page to refresh
    except (TimeoutException, NoSuchElementException):
        print("No more pages. Exiting.")
        break

# Save to CSV
if product_list:
    df = pd.DataFrame(product_list)
    df.to_csv("thriveil_products.csv", mode='a', header=not pd.io.common.file_exists("thriveil_products.csv"), index=False)
    print(f"Saved {len(product_list)} products to CSV.")
else:
    print("No data scraped.")

driver.quit()

Final Tips

  • Test selectors manually: Use your browser's dev tools (F12) to inspect elements and verify your selectors work before adding them to code.
  • Rate limiting: Add small delays (time.sleep(1-2)) between actions to avoid overwhelming the site (and getting blocked).
  • Headless mode: If you don't need to see the browser, add edge_options.add_argument("--headless=new") to run it in the background.

备注:内容来源于stack exchange,提问作者Mubaraq Onipede

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 08:24:35