使用find_elements_by_xpath匹配多位置元素时结果异常的问题
Let's break down the issues with your code and fix them step by step:
1. Why Your First XPath Only Returns the First Element
Your original XPath //li[@class='contentnode'][1 <= position() and position() < 7]/dl/dd has a subtle context ambiguity. When using // to select elements across the entire document, the position() filter can behave unexpectedly depending on the XPath implementation. A cleaner, more reliable way to target the first 6 li.contentnode elements is to simplify the filter:
Use this XPath to capture the first 6 matching li elements:
//li[@class='contentnode'][position() <= 6]/dl/dd
Or to be more precise (and avoid matching li elements from unrelated uls):
//section[@class='node_category']/ul/li[@class='contentnode'][position() <= 6]/dl/dd
2. Why the Second XPath Returns Nothing
The XPath //li[@class='contentnode'][contains(@id,'kui_3')]/dl/dd fails because not all your li.contentnode elements have an id attribute (like the second li in your HTML snippet). This filter only selects elements with an id containing "kui_3", so it automatically excludes any li without that attribute—this doesn't align with your goal of capturing all 1-6 position elements.
3. Updated Python Code
Also note that find_elements_by_xpath is deprecated in modern Selenium versions. Use the newer By.XPATH syntax instead, and add a wait to handle dynamic content (in case your elements load asynchronously):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC lst = [] browser = webdriver.Chrome('./path') url = "https://<target URL>" browser.get(url) # Wait up to 10 seconds for all target elements to load try: contents = WebDriverWait(browser, 10).until( EC.presence_of_all_elements_located( (By.XPATH, "//li[@class='contentnode'][position() <= 6]/dl/dd") ) ) for t in contents: lst.append(t.text) # Append text directly instead of wrapping in a nested list print(lst) finally: browser.quit() # Always clean up by closing the browser
Key Fixes & Notes
- The simplified
position() <= 6filter clearly targets the first 6lielements in their parent container. - Using
WebDriverWaitensures you don't try to access elements before they're fully rendered (a common issue with dynamic websites). - The deprecated
find_elements_by_xpathis replaced with the standardBy.XPATHapproach for long-term compatibility.
内容的提问来源于stack exchange,提问作者K.K.

