You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup提取嵌套标签中的指定文本?

Got it, here's a straightforward way to extract the "Find me" text from that nested blockquote structure using Python and BeautifulSoup:

First, make sure you have BeautifulSoup installed (if you don't already):

pip install beautifulsoup4

Then use this code snippet to parse and extract the text:

from bs4 import BeautifulSoup

# Paste your HTML content here
html = '''<div><blockquote type="cite" class=""><p>Find me</p> <blockquote cite="mid:609415CB-0979-47C1-9A75-CE1BE65939A0@wiwacom.fr" type="cite" class=""><p>Not me</p> <blockquote type="cite" class=""><p>Not me too</p> </blockquote> </blockquote> </div>'''

# Parse the HTML with BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')

# Option 1: Use CSS selectors to target the first (outermost) blockquote's p tag
find_me = soup.select_one('div > blockquote:first-of-type > p').get_text(strip=True)

# Option 2: Alternatively, use find() which grabs the first matching element by default
# find_me = soup.find('blockquote').find('p').get_text(strip=True)

print(find_me)  # Output: Find me

Quick breakdown:

  • select_one() uses a CSS selector to pinpoint the exact element: div > blockquote:first-of-type targets the outermost blockquote directly inside the div, then we grab its child p tag.
  • get_text(strip=True) cleans up any extra whitespace around the text so you get a neat "Find me" result.
  • The second option with find() works because soup.find('blockquote') will return the first (outermost) blockquote in the HTML, so we can just chain find('p') to get its paragraph.

内容的提问来源于stack exchange,提问作者mee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:03:39