如何使用BeautifulSoup抓取ID动态变化的<div>内的<p>标签
Hey there! Let's get your script to target only those paragraphs inside that dynamically named <div> instead of pulling every <p> from the whole page. Here's how to make it work:
Step 1: Add the regex module
First, you'll need Python's built-in re module to match that dynamic ID pattern. Make sure to import it alongside your other modules.
Step 2: Target the dynamic with regex
Instead of grabbing all <p> tags right away, first locate the <div> whose ID starts with content-body- followed by numbers. BeautifulSoup lets you use regular expressions in attribute searches, which is perfect for this dynamic ID scenario.
Full Working Code
import requests from bs4 import BeautifulSoup import re url = 'http://www.thehindu.com/opinion/op-ed/Does-Beijing-really-want-quotbreak-upquot-India/article16875298.ece' page = requests.get(url) html = page.content soup = BeautifulSoup(html, 'html.parser') # Find the div with ID matching "content-body-" + numbers target_div = soup.find('div', attrs={'id': re.compile(r'^content-body-\d+$')}) if target_div: # Now only grab <p> tags inside this specific div for tag in target_div.find_all('p'): print(tag.text.encode('utf-8').decode('utf-8') + '\n') # Get clean text instead of byte strings else: print("Couldn't find the target div with the dynamic ID!")
Quick Explanation
- The regex
r'^content-body-\d+$'ensures we only match IDs that start exactly withcontent-body-and end with one or more digits (the dynamic part). - We check if
target_divexists before looping to avoid errors if the page structure changes or the div isn't found. - Added
.decode('utf-8')so you get readable text instead of byte strings in Python 3 (skip this if you're using Python 2, but working with Unicode strings is always safer).
This way, you'll only extract paragraphs from that specific dynamic <div> instead of the entire page's content.
内容的提问来源于stack exchange,提问作者Shyam

