You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup移除class为sf-item的div内的<u>和<a>标签

Solution to Remove <u> and <a> Tags Within Specific Divs (Preserving Text)

Hey there! I get exactly what you're dealing with—trying to target only the <u> and <a> tags inside div.sf-item without nuking their text content, and avoiding accidental changes to elements outside that scope. Let's fix this properly with BeautifulSoup, since regex is almost never the right call for HTML manipulation.

The Core Issue with Your Previous Attempts

  • Regex: HTML is nested and irregular, so regex can't reliably restrict matches to only the div.sf-item context—you'll end up matching tags elsewhere or missing nested cases.
  • Blindly removing tags with BeautifulSoup: If you used decompose() or extract(), that removes the tag and its content entirely. What you need instead is to "unwrap" the tag, leaving its text behind.

Step-by-Step Implementation

Here's how to target only the desired tags within div.sf-item and preserve their text:

from bs4 import BeautifulSoup

# Sample HTML input (replace with your crawled content)
html_content = """
<div class="sf-item">
  This is <u>underlined text</u> and <a href="link">linked text</a> inside the target div.
</div>
<div>
  This <u>underlined text</u> should stay untouched outside the sf-item div.
</div>
"""

# Parse the HTML
soup = BeautifulSoup(html_content, "html.parser")

# 1. Find all divs with class "sf-item"
target_divs = soup.find_all("div", class_="sf-item")

# 2. For each target div, find all <u> and <a> tags and unwrap them
for div in target_divs:
    # Unwrap <u> tags
    for u_tag in div.find_all("u"):
        u_tag.unwrap()
    # Unwrap <a> tags
    for a_tag in div.find_all("a"):
        a_tag.unwrap()

# Get the modified HTML or extract text
modified_html = str(soup)
extracted_text = soup.get_text(strip=True, separator=" ")

print("Modified HTML:")
print(modified_html)
print("\nExtracted Text:")
print(extracted_text)

What This Does

  • unwrap() removes the tag but keeps all its child content (text, other nested tags) in place—exactly what you need to avoid splitting your crawled text.
  • By first targeting only div.sf-item, you ensure no changes are made to elements outside this scope.

Example Output

Modified HTML:

<div class="sf-item">
  This is underlined text and linked text inside the target div.
</div>
<div>
  This <u>underlined text</u> should stay untouched outside the sf-item div.
</div>

Extracted Text:

This is underlined text and linked text inside the target div. This underlined text should stay untouched outside the sf-item div.

This matches your expected outcome: the text from <u> and <a> is preserved, and only the tags inside div.sf-item are removed.

内容的提问来源于stack exchange,提问作者Satheesh Panduga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 16:57:55