如何用BeautifulSoup移除class为sf-item的div内的<u>和<a>标签
Hey there! I get exactly what you're dealing with—trying to target only the <u> and <a> tags inside div.sf-item without nuking their text content, and avoiding accidental changes to elements outside that scope. Let's fix this properly with BeautifulSoup, since regex is almost never the right call for HTML manipulation.
The Core Issue with Your Previous Attempts
- Regex: HTML is nested and irregular, so regex can't reliably restrict matches to only the
div.sf-itemcontext—you'll end up matching tags elsewhere or missing nested cases. - Blindly removing tags with BeautifulSoup: If you used
decompose()orextract(), that removes the tag and its content entirely. What you need instead is to "unwrap" the tag, leaving its text behind.
Step-by-Step Implementation
Here's how to target only the desired tags within div.sf-item and preserve their text:
from bs4 import BeautifulSoup # Sample HTML input (replace with your crawled content) html_content = """ <div class="sf-item"> This is <u>underlined text</u> and <a href="link">linked text</a> inside the target div. </div> <div> This <u>underlined text</u> should stay untouched outside the sf-item div. </div> """ # Parse the HTML soup = BeautifulSoup(html_content, "html.parser") # 1. Find all divs with class "sf-item" target_divs = soup.find_all("div", class_="sf-item") # 2. For each target div, find all <u> and <a> tags and unwrap them for div in target_divs: # Unwrap <u> tags for u_tag in div.find_all("u"): u_tag.unwrap() # Unwrap <a> tags for a_tag in div.find_all("a"): a_tag.unwrap() # Get the modified HTML or extract text modified_html = str(soup) extracted_text = soup.get_text(strip=True, separator=" ") print("Modified HTML:") print(modified_html) print("\nExtracted Text:") print(extracted_text)
What This Does
unwrap()removes the tag but keeps all its child content (text, other nested tags) in place—exactly what you need to avoid splitting your crawled text.- By first targeting only
div.sf-item, you ensure no changes are made to elements outside this scope.
Example Output
Modified HTML:
<div class="sf-item"> This is underlined text and linked text inside the target div. </div> <div> This <u>underlined text</u> should stay untouched outside the sf-item div. </div>
Extracted Text:
This is underlined text and linked text inside the target div. This underlined text should stay untouched outside the sf-item div.
This matches your expected outcome: the text from <u> and <a> is preserved, and only the tags inside div.sf-item are removed.
内容的提问来源于stack exchange,提问作者Satheesh Panduga
相关产品推荐
相关产品推荐

