You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium筛选特定字体的p元素文本(排除不符子元素)

Solution to Extract Text Matching Specific Font from <p> Elements

Great question! Your current code pulls all text inside the <p> tag, including text from child elements like <i> that don't match your target font. To fix this, we need to recursively check every node in the <p> element's hierarchy and only collect text whose parent element has the matching stored_font property.

Step 1: Create a Recursive Text Extraction Function

This function will traverse all child nodes (including raw text nodes) and filter text based on the parent element's font:

def get_matching_text(element, target_font):
    matching_text = []
    # Iterate over all child nodes of the element
    for node in element.get_property('childNodes'):
        # Handle text nodes (nodeType 3 represents TEXT_NODE)
        if node.nodeType == 3:
            # Get the parent element of this text node to check its font property
            parent_element = node.parentElement
            if parent_element.value_of_css_property('font') == target_font:
                # Clean up extra whitespace and add non-empty text to the list
                cleaned_text = node.strip()
                if cleaned_text:
                    matching_text.append(cleaned_text)
        # Handle element nodes (nodeType 1 represents ELEMENT_NODE) by recursing into them
        elif node.nodeType == 1:
            matching_text.extend(get_matching_text(node, target_font))
    # Join collected text with spaces to maintain natural sentence spacing
    return ' '.join(matching_text)

Step 2: Update Your Original Loop

Replace the direct elem.text call with our recursive function to get only the matching text:

snippets = []
for elem in br.find_elements_by_tag_name('p'):
    # First verify the <p> itself matches your target font (as you did before)
    if elem.value_of_css_property('font') == stored_font:
        # Extract only text from nodes that meet the font requirement
        matched_content = get_matching_text(elem, stored_font)
        snippets.append(matched_content)

How This Works for Your Example

For your sample <p> Hello there <i> mate </i> ! </p>:

  • The text nodes "Hello there " and " !" have the <p> element as their parent, which matches stored_font — so they get added to the result.
  • The text node "mate" has the <i> element as its parent. If the <i>'s font doesn't match stored_font, it gets excluded entirely.
  • The final output is "Hello there !", exactly what you need.

Key Notes

  • We use childNodes instead of just child elements to access raw text content directly, which is crucial for filtering at the text level.
  • Checking nodeType lets us separate plain text nodes from nested HTML elements.
  • Stripping whitespace and skipping empty text nodes ensures we don't end up with messy extra spaces in the final result.

内容的提问来源于stack exchange,提问作者greenberet123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:07:40