You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取无HTML标签的纯文本?附目标HTML代码示例

提取HTML中的纯文本方法

嘿,我来帮你搞定从这段HTML里提取纯文本的事儿!针对你给出的代码,我给你几个实用的方案:

1. 手动快速提取(适合简单HTML)

你这段HTML结构很简单,没有复杂嵌套,直接删掉所有<>包裹的标签就行,最终得到的纯文本是:

CPSC 353 - Introduction to Computer Security (3)
Security goals, security systems, access controls, networks and security, integrity, cryptography fundamentals, authentication. Attacks: software, network, website; management considerations, security standards in government and industry; security issues in requirements, architecture, design, implementation, testing, operation, maintenance, acquisition,

2. Python自动化提取(适合批量/复杂场景)

如果以后要处理大量或者结构复杂的HTML,用BeautifulSoup库绝对省心,代码示例如下:

from bs4 import BeautifulSoup

# 你的目标HTML内容
html_content = '<div class="ajaxcourseindentfix"> <h3>CPSC 353 - Introduction to Computer Security (3) </h3> <hr>Security goals, security systems, access controls, networks and security, integrity, cryptography fundamentals, authentication. Attacks: software, network, website; management considerations, security standards in government and industry; security issues in requirements, architecture, design, implementation, testing, operation, maintenance, acquisition,'

# 解析HTML结构
soup = BeautifulSoup(html_content, 'html.parser')

# 提取纯文本并整理格式
plain_text = soup.get_text(strip=True, separator=' ')
print(plain_text)

运行后会输出去掉多余空格、格式整齐的纯文本,strip=True会清理掉冗余的换行和空格,separator=' '用空格分隔不同标签里的内容。

3. JavaScript提取(前端/Node.js场景)

如果是在浏览器或者Node.js环境里处理,可以用原生的DOMParser来实现:

const htmlContent = '<div class="ajaxcourseindentfix"> <h3>CPSC 353 - Introduction to Computer Security (3) </h3> <hr>Security goals, security systems, access controls, networks and security, integrity, cryptography fundamentals, authentication. Attacks: software, network, website; management considerations, security standards in government and industry; security issues in requirements, architecture, design, implementation, testing, operation, maintenance, acquisition,';

const parser = new DOMParser();
const doc = parser.parseFromString(htmlContent, 'text/html');
const plainText = doc.body.textContent.trim();
console.log(plainText);

这个方法会自动忽略所有HTML标签,直接提取页面中的所有可见文本,非常适合前端场景。

内容的提问来源于stack exchange,提问作者miserable

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:20:22