如何用BeautifulSoup和Python requests获取含strong标签段落及下方三段内容
<strong> Tag Plus the Next 3 Paragraphs Using BeautifulSoup & Requests Hey there! Let’s sort this out for you. The main issues with your current code are:
- The target paragraph has its "Examples" text wrapped in a
<strong>tag, which can make direct text matching finicky in some scenarios. - You’re only fetching that single paragraph instead of including the three that come right after it.
Here’s a revised approach that fixes both problems:
Step 1: Reliably Locate the Target Paragraph
Instead of matching the full paragraph text with regex, we can directly check for a <p> tag that contains a <strong> element with the exact text "Examples". This is way more robust because it targets the specific element you care about.
Step 2: Grab the Target + Next 3 Paragraphs
Once we’ve found our starting paragraph, we use BeautifulSoup’s find_next_siblings() method to pull in the three subsequent <p> tags, then combine them with the original.
Here’s the full working code:
import requests from bs4 import BeautifulSoup URL = 'https://www.bbc.co.uk/learningenglish/english/features/the-english-we-speak/ep-200601' headers={ 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.61 Safari/537.36' } # Fetch and parse the webpage page = requests.get(URL, headers=headers) soup = BeautifulSoup(page.content, 'html.parser') # Find the paragraph with <strong>Examples</strong> inside it target_para = soup.find('p', lambda tag: tag.strong and tag.strong.text.strip() == 'Examples') if target_para: # Collect the target paragraph + next 3 sibling paragraphs desired_paragraphs = [target_para] desired_paragraphs.extend(target_para.find_next_siblings('p', limit=3)) # Print or process the results for num, para in enumerate(desired_paragraphs, 1): print(f"Paragraph {num}:\n{para.get_text(strip=False)}\n") else: print("Couldn't find the target paragraph!")
Quick Breakdown:
- Target Paragraph Search: The lambda function checks if a
<p>has a<strong>child, and if that child’s stripped text is exactly "Examples" — no guesswork with regex here. - Fetching Subsequent Paragraphs:
find_next_siblings('p', limit=3)grabs the next three<p>elements immediately after our target, which is exactly what you need. - Output: The loop prints each paragraph with a number for clarity, but you can easily modify this to save the text to a file or process it further.
If you prefer using regex for text matching, you could adjust your original code to use re.compile(r'^Examples') (since the paragraph’s text starts with "Examples: "), but the <strong> tag check is more reliable in case the page’s formatting changes slightly.
内容的提问来源于stack exchange,提问作者Jay Sinha

