You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将同class名的div内容爬取到数组中?

Hey there! Since you've got JavaScript experience but are new to Python and Selenium, let's walk through exactly how to scrape those dynamic product sections into a clean array (Python list, technically). Here's a step-by-step solution tailored to your use case:

Solution for Extracting Dynamic Product Sections with Selenium (Python)

Step 1: Get Your Tools Ready

First, make sure you have the necessary packages installed and drivers set up:

  • Install Selenium via pip:
    pip install selenium
    
  • Download the appropriate browser driver (e.g., ChromeDriver for Google Chrome) and ensure it's in your system PATH or specify its path directly in the code.

Step 2: Full Code Implementation

This script will wait for the dynamic content to load, loop through each product section, extract the title and associated products, and store everything in a list of dictionaries (which acts like your "array" of section data):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize the Chrome driver (adjust if using Firefox/Edge)
driver = webdriver.Chrome()

try:
    # Navigate to your target page
    driver.get("YOUR_TARGET_URL_HERE")

    # Wait for the first product section to load (adjust timeout as needed)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "product-section"))
    )

    # Initialize an empty list to store all section data
    product_sections = []

    # Get all product sections on the page
    sections = driver.find_elements(By.CLASS_NAME, "product-section")

    for section in sections:
        # Extract the section title
        section_title = section.find_element(By.CLASS_NAME, "section-title").text.strip()

        # Extract all products in this section
        product_list = section.find_element(By.CLASS_NAME, "products")
        product_items = product_list.find_elements(By.TAG_NAME, "li")
        
        products = []
        for item in product_items:
            # Example: Extract product name and URL (adjust based on actual element structure)
            product_name = item.find_element(By.TAG_NAME, "h3").text.strip()
            product_url = item.find_element(By.TAG_NAME, "a").get_attribute("href")
            products.append({
                "name": product_name,
                "url": product_url
            })

        # Add the section data to our main list
        product_sections.append({
            "section_title": section_title,
            "products": products
        })

    # Print the result to verify (you can save this to a file later)
    print(product_sections)

finally:
    # Always close the driver when done
    driver.quit()

Key Explanations for JavaScript Devs

  • WebDriverWait: Instead of using time.sleep() (which is unreliable), this waits until the critical elements are present—similar to using await with DOM queries in JS.
  • Element Locators: By.CLASS_NAME, By.TAG_NAME are like using document.getElementsByClassName() or document.getElementsByTagName() in JS.
  • Python Lists/Dicts: The product_sections list is your "array", and each section is a dictionary (like a JS object) holding the title and product data.

Troubleshooting Tips

  • If elements aren't found: Double-check the class names (they're case-sensitive!) or use more specific locators (e.g., By.CSS_SELECTOR if classes are combined).
  • Slow loading pages: Increase the timeout in WebDriverWait(driver, 10) (change 10 to 15/20 seconds).
  • Dynamic content that loads on scroll: You might need to add code to scroll the page and wait for more sections to load—feel free to ask for help with that if needed!

内容的提问来源于stack exchange,提问作者user2868900

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:40:37