You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup根据h2文本提取指定段落数据?

Solution to Extract Target H2 and Corresponding Paragraphs

Got it, let's fix this up for you. Your current code only pulls all the h2 text, but we need two key things here: link each h2 to its following paragraph and filter only the headings you care about (Test1 and Test3).

Here's the modified code that does exactly what you need:

from bs4 import BeautifulSoup

# Replace this with your actual page.text content
page_text = '''<div class="myClass"> <div itemprop="reviewBody" class="review-body"> <h2 class="h3">Test1</h2><p>I want to extract this</p> <h2 class="h3">Test2</h2><p>Dont want to extract</p> <h2 class="h3">Test3</h2><p>I want to extract this too</p> </div> </div>'''

soup = BeautifulSoup(page_text, 'html.parser')
# Get the parent div (no need for find_all since we only need the first one)
target_div = soup.find(class_="myClass")

# Define the headings you want to keep - use a set for fast lookups
wanted_headings = {"Test1", "Test3"}

# Iterate through each h2 in the target div
for h2 in target_div.find_all('h2', class_="h3"):
    heading_text = h2.text.strip()
    # Check if this heading is in our wanted list
    if heading_text in wanted_headings:
        # Grab the immediately following paragraph using find_next()
        para_text = h2.find_next('p').text.strip()
        # Print in your desired format
        print(f"{heading_text} | {para_text}")

Key Changes Explained:

  • Filtering Target Headings: We use a set wanted_headings to quickly check if the current h2 is one we want to extract. Sets are faster than lists for membership checks, which is helpful if you have more headings later.
  • Linking H2 to Paragraph: The find_next('p') method grabs the first <p> tag that comes right after the h2, ensuring we get the correct corresponding paragraph.
  • Simplified Parent Selection: Using soup.find(class_="myClass") instead of find_all()[0] is cleaner and more readable when you only need the first matching div.

When you run this code, you'll get exactly the output you want:

Test1 | I want to extract this
Test3 | I want to extract this too

内容的提问来源于stack exchange,提问作者codecodecode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:54:58