使用Selenium抓取Thriveil网站产品信息时元素检测失败的求助
Hey there! I took a close look at your code and the problem you're facing with element detection on the Thriveil site. Let's break down the key issues and fix them step by step—no jargon, just practical changes to get your scraper working reliably.
1. Critical Bug: Incorrect Data Collection Logic
Right now, your code is appending entire lists (like names, brand_names) to each product entry, which is not only wrong but can also cause unexpected behavior. Instead, you should collect individual product data for each item and add that single entry to your product_list.
For example, instead of:
product_data = { "name": names, "brand_name": brand_names, ... } product_list.append(product_data)
You should do:
# For each product, create a single dict with its data product_data = { "name": name.text.strip(), "brand_name": brand.text.strip(), "brand_link": brand_element.get_attribute('href'), "strain": strain.text.strip() if strain else "N/A", "potency": potency_text if 'potency_text' in locals() else "N/A", "price": price.text.strip() if price else "N/A", "effects": ", ".join([e.text.strip() for e in effect_elements]) if effect_elements else "N/A" } product_list.append(product_data)
And you can remove the global lists (names, brand_names, etc.) entirely—they're not needed anymore.
2. Use Targeted Waits for Every Element
Waiting just for the <h1> isn't enough. Many elements load asynchronously after the page initializes. For each element you want to scrape, use WebDriverWait to ensure it's present before trying to access it. This will fix most "element not found" errors.
Example for scraping the product name:
# Wait for the product name to be present name = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "h1[data-testid='product-name']"))).text.strip()
Do this for every element (brand, strain, price, etc.) instead of using driver.find_element directly.
3. Avoid Fragile Auto-Generated Class Names
Class names like full-card_Wrapper-sc-11z5u35-0 or typography__Brand-sc-1q7gvs8-2.fyoohd are auto-generated by frameworks like React/Vue and can change anytime the site updates. Instead, use more stable selectors:
- Data attributes: The site uses
data-testid(likedata-testid='product-name') which is designed for testing/scraping and rarely changes. - Nested selectors: If data attributes aren't available, use parent-child relationships (e.g.,
div.product-brand ainstead of a random class).
For the brand element, try this instead of the fragile class:
brand_element = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='Brand'] a"))) brand_name = brand_element.text.strip() brand_link = brand_element.get_attribute('href')
4. Bypass Selenium Detection
Many modern sites block Selenium by detecting its unique browser properties. To fix this, add stealth settings to your Edge options to make your scraper look like a regular user:
edge_options = Options() # Add anti-detection flags edge_options.add_argument("--disable-blink-features=AutomationControlled") edge_options.add_experimental_option("excludeSwitches", ["enable-automation"]) edge_options.add_experimental_option('useAutomationExtension', False) # Set a real user-agent (you can get yours from your browser's dev tools) edge_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36 Edg/120.0.0.0")
This will help the site not recognize you're using Selenium.
5. Fix Potency Extraction (Avoid Index Errors)
Your code uses potency_values[1] which will crash if there are fewer than 2 potency elements. Add a check to handle this:
potency_text = "N/A" potency_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.info-chip__InfoChipText-sc-11n9ujc-0"))) if len(potency_elements) >=2: potency_text = potency_elements[1].text.split(":")[-1].strip()
Revised Code Snippet (Key Fixes Included)
Here's a trimmed version of your code with all the above fixes applied:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.edge.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.edge.options import Options from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, TimeoutException import pandas as pd import time # Set up Edge with anti-detection edge_options = Options() edge_options.add_argument("--disable-blink-features=AutomationControlled") edge_options.add_experimental_option("excludeSwitches", ["enable-automation"]) edge_options.add_experimental_option('useAutomationExtension', False) edge_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36 Edg/120.0.0.0") service = Service("C:\\Users\\iyush\\Documents\\VS Code\\Selenium\\msedgedriver.exe") driver = webdriver.Edge(service=service, options=edge_options) url = "https://thriveil.com/casey-rec-menu/?dtche%5Bpath%5D=products" driver.get(url) wait = WebDriverWait(driver, 30) # Close ad and cookies try: close_button = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "terpli-close"))) close_button.click() except TimeoutException: print("Ad close button not found. Continuing...") try: accept_button = wait.until(EC.element_to_be_clickable((By.ID, "wt-cli-accept-all-btn"))) accept_button.click() except TimeoutException: print("Cookie consent button not found. Continuing...") product_list = [] current_page = 1 while True: # Wait for product cards to load wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div[class*='full-card_Wrapper']"))) product_cards = driver.find_elements(By.CSS_SELECTOR, "div[class*='full-card_Wrapper'] a") product_urls = [card.get_attribute("href") for card in product_cards if card.get_attribute("href")] for url in product_urls: driver.execute_script(f"window.open('{url}', '_blank');") driver.switch_to.window(driver.window_handles[-1]) try: # Wait for critical elements to load wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "h1[data-testid='product-name']"))) # Scrape each field with waits name = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "h1[data-testid='product-name']"))).text.strip() brand_element = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='Brand'] a"))) brand_name = brand_element.text.strip() brand_link = brand_element.get_attribute("href") strain = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "span[data-testid='info-chip']"))).text.strip() # Potency handling potency_text = "N/A" potency_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.info-chip__InfoChipText-sc-11n9ujc-0"))) if len(potency_elements) >=2: potency_text = potency_elements[1].text.split(":")[-1].strip() price = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='price__PriceText']"))).text.strip() effects = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.effect-tile__Text-sc-1as4rkm-1"))) effect_text = ", ".join([e.text.strip() for e in effects]) if effects else "N/A" # Add single product data to list product_list.append({ "name": name, "brand_name": brand_name, "brand_link": brand_link, "strain": strain, "potency": potency_text, "price": price, "effects": effect_text }) except Exception as e: print(f"Error scraping {url}: {str(e)}") finally: driver.close() driver.switch_to.window(driver.window_handles[0]) print(f"Page {current_page} scraped successfully.") # Next page handling try: next_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button[aria-label*='next page']"))) # Check if next button is disabled if "disabled" in next_btn.get_attribute("class"): break next_btn.click() current_page +=1 time.sleep(2) # Short wait for page to refresh except (TimeoutException, NoSuchElementException): print("No more pages. Exiting.") break # Save to CSV if product_list: df = pd.DataFrame(product_list) df.to_csv("thriveil_products.csv", mode='a', header=not pd.io.common.file_exists("thriveil_products.csv"), index=False) print(f"Saved {len(product_list)} products to CSV.") else: print("No data scraped.") driver.quit()
Final Tips
- Test selectors manually: Use your browser's dev tools (F12) to inspect elements and verify your selectors work before adding them to code.
- Rate limiting: Add small delays (
time.sleep(1-2)) between actions to avoid overwhelming the site (and getting blocked). - Headless mode: If you don't need to see the browser, add
edge_options.add_argument("--headless=new")to run it in the background.
备注:内容来源于stack exchange,提问作者Mubaraq Onipede

