如何提取无HTML标签的纯文本?附目标HTML代码示例
嘿,我来帮你搞定从这段HTML里提取纯文本的事儿!针对你给出的代码,我给你几个实用的方案:
1. 手动快速提取(适合简单HTML)
你这段HTML结构很简单,没有复杂嵌套,直接删掉所有<>包裹的标签就行,最终得到的纯文本是:
CPSC 353 - Introduction to Computer Security (3)
Security goals, security systems, access controls, networks and security, integrity, cryptography fundamentals, authentication. Attacks: software, network, website; management considerations, security standards in government and industry; security issues in requirements, architecture, design, implementation, testing, operation, maintenance, acquisition,
2. Python自动化提取(适合批量/复杂场景)
如果以后要处理大量或者结构复杂的HTML,用BeautifulSoup库绝对省心,代码示例如下:
from bs4 import BeautifulSoup # 你的目标HTML内容 html_content = '<div class="ajaxcourseindentfix"> <h3>CPSC 353 - Introduction to Computer Security (3) </h3> <hr>Security goals, security systems, access controls, networks and security, integrity, cryptography fundamentals, authentication. Attacks: software, network, website; management considerations, security standards in government and industry; security issues in requirements, architecture, design, implementation, testing, operation, maintenance, acquisition,' # 解析HTML结构 soup = BeautifulSoup(html_content, 'html.parser') # 提取纯文本并整理格式 plain_text = soup.get_text(strip=True, separator=' ') print(plain_text)
运行后会输出去掉多余空格、格式整齐的纯文本,strip=True会清理掉冗余的换行和空格,separator=' '用空格分隔不同标签里的内容。
3. JavaScript提取(前端/Node.js场景)
如果是在浏览器或者Node.js环境里处理,可以用原生的DOMParser来实现:
const htmlContent = '<div class="ajaxcourseindentfix"> <h3>CPSC 353 - Introduction to Computer Security (3) </h3> <hr>Security goals, security systems, access controls, networks and security, integrity, cryptography fundamentals, authentication. Attacks: software, network, website; management considerations, security standards in government and industry; security issues in requirements, architecture, design, implementation, testing, operation, maintenance, acquisition,'; const parser = new DOMParser(); const doc = parser.parseFromString(htmlContent, 'text/html'); const plainText = doc.body.textContent.trim(); console.log(plainText);
这个方法会自动忽略所有HTML标签,直接提取页面中的所有可见文本,非常适合前端场景。
内容的提问来源于stack exchange,提问作者miserable

