如何用Python与Selenium提取dd标签内价格并去除美元符号?
Fixing Selenium Extraction for dd Tag Text
Let’s walk through fixing your code and addressing both scenarios (your specific dd tag with attributes, and dd tags without named attributes):
Issues in Your Current Code
find_elements_by_xpathreturns a list of elements, not a single element. Callingget_attributeon a list will throw an error—usefind_element_by_xpath(singular) instead if targeting one specific element.- The price you want is the text content of the dd tag, not an attribute value. Use the
.textproperty instead ofget_attribute(). - Your XPath is too broad (
//dl[@class='dl']//dd) which might match multiple elements. Let’s narrow it down to target the exact dd you need.
Solution for Your Specific dd Tag
Your dd tag has a unique itemprop="youSave" attribute—we can use that to locate it directly. Here’s the corrected code, including waiting for the element to load (to avoid race conditions with dynamic content):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize driver (adjust based on your browser) driver = webdriver.Chrome() driver.get("your_page_url_here") try: # Wait up to 10 seconds for the element to be visible price_saved_element = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//dd[@itemprop='youSave']")) ) # Get the text and remove the dollar sign price_saved_text = price_saved_element.text.strip() price_without_dollar = price_saved_text.replace("$", "") print(f"Saved price (without $): {price_without_dollar}") except Exception as e: print(f"Failed to extract price: {str(e)}") finally: driver.quit()
Extracting Text from dd Tags Without Named Attributes
If you have a dd tag with no unique attributes (no itemprop, class, or id), you can locate it using its position relative to parent elements or sibling tags:
- If it’s the 2nd dd inside a dl with class
dl:dd_element = driver.find_element(By.XPATH, "//dl[@class='dl']//dd[2]") dd_text = dd_element.text.strip() - If it’s immediately after a specific dt tag (e.g., a dt with text "You Save"):
dd_element = driver.find_element(By.XPATH, "//dt[text()='You Save']/following-sibling::dd[1]") dd_text = dd_element.text.strip()
These methods let you target even unlabeled dd tags by their context in the HTML structure.
内容的提问来源于stack exchange,提问作者Bob Stone
相关产品推荐
相关产品推荐

