You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何下载去除所有类、ID、属性的网站纯HTML代码?

如何去除HTML中所有class、id及属性内容

下面提供几种实用的实现方式:

方法1:用Python的BeautifulSoup库处理(推荐,兼容复杂HTML)

BeautifulSoup是处理HTML的专业库,能准确解析标签结构,避免正则的局限性。

首先安装库:

pip install beautifulsoup4

然后运行以下代码:

from bs4 import BeautifulSoup

# 替换成你的原始HTML内容
raw_html = """
<h1 id='page-title'>This is a page</h1>
<div class='page-part'>
    <button id='red-button' style='background-color:Red'>I'm a button</button>
    <button id='blue-button' style='background-color:Blue'>I'm a button</button>
</div>
"""

# 解析HTML
soup = BeautifulSoup(raw_html, 'html.parser')

# 遍历所有标签,清空属性
for tag in soup.find_all():
    tag.attrs = {}

# 输出格式化后的纯结构HTML
print(soup.prettify())

运行后就能得到仅保留标签结构和文本内容的HTML。

方法2:正则表达式快速处理(适合简单HTML场景)

如果你的HTML结构简单,没有复杂嵌套或特殊字符属性,可以用正则快速替换:

import re

raw_html = """
<h1 id='page-title'>This is a page</h1>
<div class='page-part'>
    <button id='red-button' style='background-color:Red'>I'm a button</button>
    <button id='blue-button' style='background-color:Blue'>I'm a button</button>
</div>
"""

# 匹配标签开头的属性部分并移除
cleaned_html = re.sub(r'<(\w+)\s+[^>]+>', r'<\1>', raw_html)
print(cleaned_html)

注意:正则无法处理属性中包含>、嵌套引号等复杂情况,若HTML结构复杂,优先用BeautifulSoup。

方法3:浏览器控制台临时处理网页

如果只是想快速处理当前打开的网页,按F12打开浏览器控制台,执行以下JavaScript代码:

// 遍历页面所有元素,移除所有属性
document.querySelectorAll('*').forEach(el => {
  Array.from(el.attributes).forEach(attr => el.removeAttribute(attr.name));
});

// 打印处理后的完整HTML,也可以直接右键页面复制
console.log(document.documentElement.outerHTML);

执行后,页面的所有元素属性都会被清除,此时复制页面HTML即可得到目标内容。

内容的提问来源于stack exchange,提问作者MehranJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 11:25:57