如何用Python将同class名的div内容爬取到数组中?
Hey there! Since you've got JavaScript experience but are new to Python and Selenium, let's walk through exactly how to scrape those dynamic product sections into a clean array (Python list, technically). Here's a step-by-step solution tailored to your use case:
Solution for Extracting Dynamic Product Sections with Selenium (Python)
Step 1: Get Your Tools Ready
First, make sure you have the necessary packages installed and drivers set up:
- Install Selenium via pip:
pip install selenium - Download the appropriate browser driver (e.g., ChromeDriver for Google Chrome) and ensure it's in your system PATH or specify its path directly in the code.
Step 2: Full Code Implementation
This script will wait for the dynamic content to load, loop through each product section, extract the title and associated products, and store everything in a list of dictionaries (which acts like your "array" of section data):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize the Chrome driver (adjust if using Firefox/Edge) driver = webdriver.Chrome() try: # Navigate to your target page driver.get("YOUR_TARGET_URL_HERE") # Wait for the first product section to load (adjust timeout as needed) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "product-section")) ) # Initialize an empty list to store all section data product_sections = [] # Get all product sections on the page sections = driver.find_elements(By.CLASS_NAME, "product-section") for section in sections: # Extract the section title section_title = section.find_element(By.CLASS_NAME, "section-title").text.strip() # Extract all products in this section product_list = section.find_element(By.CLASS_NAME, "products") product_items = product_list.find_elements(By.TAG_NAME, "li") products = [] for item in product_items: # Example: Extract product name and URL (adjust based on actual element structure) product_name = item.find_element(By.TAG_NAME, "h3").text.strip() product_url = item.find_element(By.TAG_NAME, "a").get_attribute("href") products.append({ "name": product_name, "url": product_url }) # Add the section data to our main list product_sections.append({ "section_title": section_title, "products": products }) # Print the result to verify (you can save this to a file later) print(product_sections) finally: # Always close the driver when done driver.quit()
Key Explanations for JavaScript Devs
- WebDriverWait: Instead of using
time.sleep()(which is unreliable), this waits until the critical elements are present—similar to usingawaitwith DOM queries in JS. - Element Locators:
By.CLASS_NAME,By.TAG_NAMEare like usingdocument.getElementsByClassName()ordocument.getElementsByTagName()in JS. - Python Lists/Dicts: The
product_sectionslist is your "array", and each section is a dictionary (like a JS object) holding the title and product data.
Troubleshooting Tips
- If elements aren't found: Double-check the class names (they're case-sensitive!) or use more specific locators (e.g.,
By.CSS_SELECTORif classes are combined). - Slow loading pages: Increase the timeout in
WebDriverWait(driver, 10)(change 10 to 15/20 seconds). - Dynamic content that loads on scroll: You might need to add code to scroll the page and wait for more sections to load—feel free to ask for help with that if needed!
内容的提问来源于stack exchange,提问作者user2868900
相关产品推荐
相关产品推荐

