如何用BeautifulSoup获取指定节点后的div/p节点并遇h1/2/3停止
Solution to Collect Div/P Nodes Until H1/H2/H3
Got it, let's tackle this problem step by step. You've already nailed getting that target red-highlighted node with your find call, so now we just need to gather all the following <div> and <p> nodes until we hit an <h1>, <h2>, or <h3>.
Approach for Sibling Nodes (Most Likely Scenario)
If the blue-marked nodes are siblings that come right after your target node, use this sibling traversal method:
# Your existing target node is stored in the `title` variable result_nodes = [] current_node = title.next_sibling while current_node is not None: # Skip empty text nodes (like newlines/spaces in raw HTML) if current_node.name is None: current_node = current_node.next_sibling continue # Stop immediately when we hit any heading tag if current_node.name in ['h1', 'h2', 'h3']: break # Add valid div/p nodes to our result list if current_node.name in ['div', 'p']: result_nodes.append(current_node) # Move to the next sibling node current_node = current_node.next_sibling
How This Works:
- We start traversing right after your target node using
next_sibling - We filter out blank text nodes (BeautifulSoup treats whitespace/newlines as separate nodes, so we need to skip those to avoid clutter)
- As soon as we encounter an h1/h2/h3, we break the loop to stop collecting further nodes
- Any div or p nodes found before hitting those headings get added to our result list
If Target Nodes Are Direct Children
If the blue-marked nodes are direct children of your red-highlighted target node instead of siblings, adjust the code to iterate over the target's children:
result_nodes = [] for child in title.children: if child.name is None: continue if child.name in ['h1', 'h2', 'h3']: break if child.name in ['div', 'p']: result_nodes.append(child)
Either implementation will give you exactly the list of nodes you need, stopping as soon as those heading tags are encountered.
内容的提问来源于stack exchange,提问作者user8162574
相关产品推荐
相关产品推荐

