如何用XPath提取HTML DOM文本节点?解决Web of Science分类获取难题
Let's break down why your current code isn't working, then walk through a few solid solutions to grab those category labels like "Mathematics" or "Applied Mathematics".
Why Your Original Code Fails
Your XPath //*[contains(text(), 'Web of Science Categories:')] is targeting the <span class="FR_label"> element itself, not the parent container that holds both the label and the category text. That's why .text only returns "Web of Science Categories:"—the actual category is a sibling text node right after the span, not part of the span's content.
Solution 1: Grab the Parent Element's Text & Clean It Up
The simplest approach is to target the parent <p class="FR_field"> element, get its full text, then strip out the label prefix:
# Locate the parent p element that contains the category label span category_container = driver.find_element_by_xpath("//p[contains(span/text(), 'Web of Science Categories:')]") # Extract full text and remove the label, then trim extra whitespace category_text = category_container.text.replace("Web of Science Categories:", "").strip()
This works because the parent p element's .text property combines all text inside it (both the span's label and the category text). The strip() handles any extra spaces around the category.
Solution 2: Directly Target the Sibling Text Node with JavaScript
If you want to be more precise (and avoid relying on text replacement), use JavaScript to fetch the text node immediately after the label span. This bypasses any issues with dynamic container elements:
category_text = driver.execute_script(""" // Find the label span, then get its next sibling text node const labelSpan = document.querySelector('span.FR_label:contains("Web of Science Categories:")'); return labelSpan ? labelSpan.nextSibling.textContent.trim() : ''; """)
Or use XPath in the JavaScript execution for even more flexibility:
category_text = driver.execute_script(""" return document.evaluate( '//span[contains(text(), "Web of Science Categories:")]/following-sibling::text()[1]', document, null, XPathResult.STRING_TYPE, null ).stringValue.trim(); """)
Solution 3: Parse with BeautifulSoup (Great for Complex Pages)
If you're dealing with lots of dynamic content or need to extract multiple categories, pairing Selenium with BeautifulSoup makes parsing easier:
from bs4 import BeautifulSoup # Get the full page source from Selenium page_html = driver.page_source soup = BeautifulSoup(page_html, 'html.parser') # Find the exact label span category_label = soup.find('span', class_='FR_label', text='Web of Science Categories:') if category_label: # Grab the text right after the span and clean it up category_text = category_label.next_sibling.strip()
This method is especially useful if categories are listed in a comma-separated format—you can split the text into a list with category_text.split(', ') to get individual categories.
内容的提问来源于stack exchange,提问作者Aaron Cao

