You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup和Python requests获取含strong标签段落及下方三段内容

How to Get a Paragraph with <strong> Tag Plus the Next 3 Paragraphs Using BeautifulSoup & Requests

Hey there! Let’s sort this out for you. The main issues with your current code are:

  • The target paragraph has its "Examples" text wrapped in a <strong> tag, which can make direct text matching finicky in some scenarios.
  • You’re only fetching that single paragraph instead of including the three that come right after it.

Here’s a revised approach that fixes both problems:

Step 1: Reliably Locate the Target Paragraph

Instead of matching the full paragraph text with regex, we can directly check for a <p> tag that contains a <strong> element with the exact text "Examples". This is way more robust because it targets the specific element you care about.

Step 2: Grab the Target + Next 3 Paragraphs

Once we’ve found our starting paragraph, we use BeautifulSoup’s find_next_siblings() method to pull in the three subsequent <p> tags, then combine them with the original.

Here’s the full working code:

import requests
from bs4 import BeautifulSoup

URL = 'https://www.bbc.co.uk/learningenglish/english/features/the-english-we-speak/ep-200601'
headers={
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.61 Safari/537.36'
}

# Fetch and parse the webpage
page = requests.get(URL, headers=headers)
soup = BeautifulSoup(page.content, 'html.parser')

# Find the paragraph with <strong>Examples</strong> inside it
target_para = soup.find('p', lambda tag: tag.strong and tag.strong.text.strip() == 'Examples')

if target_para:
    # Collect the target paragraph + next 3 sibling paragraphs
    desired_paragraphs = [target_para]
    desired_paragraphs.extend(target_para.find_next_siblings('p', limit=3))
    
    # Print or process the results
    for num, para in enumerate(desired_paragraphs, 1):
        print(f"Paragraph {num}:\n{para.get_text(strip=False)}\n")
else:
    print("Couldn't find the target paragraph!")

Quick Breakdown:

  • Target Paragraph Search: The lambda function checks if a <p> has a <strong> child, and if that child’s stripped text is exactly "Examples" — no guesswork with regex here.
  • Fetching Subsequent Paragraphs: find_next_siblings('p', limit=3) grabs the next three <p> elements immediately after our target, which is exactly what you need.
  • Output: The loop prints each paragraph with a number for clarity, but you can easily modify this to save the text to a file or process it further.

If you prefer using regex for text matching, you could adjust your original code to use re.compile(r'^Examples') (since the paragraph’s text starts with "Examples: "), but the <strong> tag check is more reliable in case the page’s formatting changes slightly.

内容的提问来源于stack exchange,提问作者Jay Sinha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:02:29