You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取DIV类H标签中strong外的指定文本

Got it! To extract just the part after "Brookstone" from that <h2> tag, you can leverage BeautifulSoup's ability to navigate sibling nodes in the HTML structure. Here's how to modify your code:

# First, grab the h2 element inside the content div
h2_element = soup.find('div', attrs={'class':'content'}).h2
# Access the text node immediately following the <strong> tag, then clean up whitespace
target_text = h2_element.strong.next_sibling.strip()
print(target_text)  # Output: AS20194 Multi-functional Massage Chair

Why this works:

The <strong>Brookstone</strong> is a direct child of the <h2> tag. The text you want exists as a plain text node (a NavigableString in BeautifulSoup terms) that sits right next to the <strong> element. Using .next_sibling targets that adjacent text node, and .strip() removes any leading/trailing spaces or line breaks that might be present in the raw HTML.

Alternative (text splitting method):

If you’re confident the "Brookstone " prefix will always be consistent, you can split the full text string instead:

full_text = soup.find('div', attrs={'class':'content'}).h2.text
target_text = full_text.split('Brookstone ', 1)[1]

This splits the full text once at "Brookstone " and takes the second segment of the resulting list. Note that this is less robust than the sibling navigation method if the prefix ever changes (e.g., extra spaces, capitalization shifts).

内容的提问来源于stack exchange,提问作者PURWU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:07:28