如何用Selenium筛选特定字体的p元素文本(排除不符子元素)
Solution to Extract Text Matching Specific Font from
<p> Elements Great question! Your current code pulls all text inside the <p> tag, including text from child elements like <i> that don't match your target font. To fix this, we need to recursively check every node in the <p> element's hierarchy and only collect text whose parent element has the matching stored_font property.
Step 1: Create a Recursive Text Extraction Function
This function will traverse all child nodes (including raw text nodes) and filter text based on the parent element's font:
def get_matching_text(element, target_font): matching_text = [] # Iterate over all child nodes of the element for node in element.get_property('childNodes'): # Handle text nodes (nodeType 3 represents TEXT_NODE) if node.nodeType == 3: # Get the parent element of this text node to check its font property parent_element = node.parentElement if parent_element.value_of_css_property('font') == target_font: # Clean up extra whitespace and add non-empty text to the list cleaned_text = node.strip() if cleaned_text: matching_text.append(cleaned_text) # Handle element nodes (nodeType 1 represents ELEMENT_NODE) by recursing into them elif node.nodeType == 1: matching_text.extend(get_matching_text(node, target_font)) # Join collected text with spaces to maintain natural sentence spacing return ' '.join(matching_text)
Step 2: Update Your Original Loop
Replace the direct elem.text call with our recursive function to get only the matching text:
snippets = [] for elem in br.find_elements_by_tag_name('p'): # First verify the <p> itself matches your target font (as you did before) if elem.value_of_css_property('font') == stored_font: # Extract only text from nodes that meet the font requirement matched_content = get_matching_text(elem, stored_font) snippets.append(matched_content)
How This Works for Your Example
For your sample <p> Hello there <i> mate </i> ! </p>:
- The text nodes
"Hello there "and" !"have the<p>element as their parent, which matchesstored_font— so they get added to the result. - The text node
"mate"has the<i>element as its parent. If the<i>'s font doesn't matchstored_font, it gets excluded entirely. - The final output is
"Hello there !", exactly what you need.
Key Notes
- We use
childNodesinstead of just child elements to access raw text content directly, which is crucial for filtering at the text level. - Checking
nodeTypelets us separate plain text nodes from nested HTML elements. - Stripping whitespace and skipping empty text nodes ensures we don't end up with messy extra spaces in the final result.
内容的提问来源于stack exchange,提问作者greenberet123
相关产品推荐
相关产品推荐

