You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XPath提取HTML DOM文本节点?解决Web of Science分类获取难题

Fix: Extracting Web of Science Category Text with Dynamic XPaths

Let's break down why your current code isn't working, then walk through a few solid solutions to grab those category labels like "Mathematics" or "Applied Mathematics".

Why Your Original Code Fails

Your XPath //*[contains(text(), 'Web of Science Categories:')] is targeting the <span class="FR_label"> element itself, not the parent container that holds both the label and the category text. That's why .text only returns "Web of Science Categories:"—the actual category is a sibling text node right after the span, not part of the span's content.

Solution 1: Grab the Parent Element's Text & Clean It Up

The simplest approach is to target the parent <p class="FR_field"> element, get its full text, then strip out the label prefix:

# Locate the parent p element that contains the category label span
category_container = driver.find_element_by_xpath("//p[contains(span/text(), 'Web of Science Categories:')]")
# Extract full text and remove the label, then trim extra whitespace
category_text = category_container.text.replace("Web of Science Categories:", "").strip()

This works because the parent p element's .text property combines all text inside it (both the span's label and the category text). The strip() handles any extra spaces around the category.

Solution 2: Directly Target the Sibling Text Node with JavaScript

If you want to be more precise (and avoid relying on text replacement), use JavaScript to fetch the text node immediately after the label span. This bypasses any issues with dynamic container elements:

category_text = driver.execute_script("""
// Find the label span, then get its next sibling text node
const labelSpan = document.querySelector('span.FR_label:contains("Web of Science Categories:")');
return labelSpan ? labelSpan.nextSibling.textContent.trim() : '';
""")

Or use XPath in the JavaScript execution for even more flexibility:

category_text = driver.execute_script("""
return document.evaluate(
  '//span[contains(text(), "Web of Science Categories:")]/following-sibling::text()[1]',
  document,
  null,
  XPathResult.STRING_TYPE,
  null
).stringValue.trim();
""")

Solution 3: Parse with BeautifulSoup (Great for Complex Pages)

If you're dealing with lots of dynamic content or need to extract multiple categories, pairing Selenium with BeautifulSoup makes parsing easier:

from bs4 import BeautifulSoup

# Get the full page source from Selenium
page_html = driver.page_source
soup = BeautifulSoup(page_html, 'html.parser')

# Find the exact label span
category_label = soup.find('span', class_='FR_label', text='Web of Science Categories:')
if category_label:
    # Grab the text right after the span and clean it up
    category_text = category_label.next_sibling.strip()

This method is especially useful if categories are listed in a comma-separated format—you can split the text into a list with category_text.split(', ') to get individual categories.


内容的提问来源于stack exchange,提问作者Aaron Cao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:23:33