如何用Python的BeautifulSoup提取嵌套标签中的指定文本?
Got it, here's a straightforward way to extract the "Find me" text from that nested blockquote structure using Python and BeautifulSoup:
First, make sure you have BeautifulSoup installed (if you don't already):
pip install beautifulsoup4
Then use this code snippet to parse and extract the text:
from bs4 import BeautifulSoup # Paste your HTML content here html = '''<div><blockquote type="cite" class=""><p>Find me</p> <blockquote cite="mid:609415CB-0979-47C1-9A75-CE1BE65939A0@wiwacom.fr" type="cite" class=""><p>Not me</p> <blockquote type="cite" class=""><p>Not me too</p> </blockquote> </blockquote> </div>''' # Parse the HTML with BeautifulSoup soup = BeautifulSoup(html, 'html.parser') # Option 1: Use CSS selectors to target the first (outermost) blockquote's p tag find_me = soup.select_one('div > blockquote:first-of-type > p').get_text(strip=True) # Option 2: Alternatively, use find() which grabs the first matching element by default # find_me = soup.find('blockquote').find('p').get_text(strip=True) print(find_me) # Output: Find me
Quick breakdown:
select_one()uses a CSS selector to pinpoint the exact element:div > blockquote:first-of-typetargets the outermost blockquote directly inside the div, then we grab its childptag.get_text(strip=True)cleans up any extra whitespace around the text so you get a neat "Find me" result.- The second option with
find()works becausesoup.find('blockquote')will return the first (outermost) blockquote in the HTML, so we can just chainfind('p')to get its paragraph.
内容的提问来源于stack exchange,提问作者mee
相关产品推荐
相关产品推荐

