如何使用BeautifulSoup提取DIV类H标签中strong外的指定文本
Got it! To extract just the part after "Brookstone" from that <h2> tag, you can leverage BeautifulSoup's ability to navigate sibling nodes in the HTML structure. Here's how to modify your code:
# First, grab the h2 element inside the content div h2_element = soup.find('div', attrs={'class':'content'}).h2 # Access the text node immediately following the <strong> tag, then clean up whitespace target_text = h2_element.strong.next_sibling.strip() print(target_text) # Output: AS20194 Multi-functional Massage Chair
Why this works:
The <strong>Brookstone</strong> is a direct child of the <h2> tag. The text you want exists as a plain text node (a NavigableString in BeautifulSoup terms) that sits right next to the <strong> element. Using .next_sibling targets that adjacent text node, and .strip() removes any leading/trailing spaces or line breaks that might be present in the raw HTML.
Alternative (text splitting method):
If you’re confident the "Brookstone " prefix will always be consistent, you can split the full text string instead:
full_text = soup.find('div', attrs={'class':'content'}).h2.text target_text = full_text.split('Brookstone ', 1)[1]
This splits the full text once at "Brookstone " and takes the second segment of the resulting list. Note that this is less robust than the sibling navigation method if the prefix ever changes (e.g., extra spaces, capitalization shifts).
内容的提问来源于stack exchange,提问作者PURWU

