如何用Python和BeautifulSoup将<p>标签替换为换行符?
移除
标签并替换为换行的BeautifulSoup实现
你可以通过两种简单方法实现需求:
方法一:使用get_text()的分隔符参数
直接调用元素的get_text()方法并指定separator='\n',就能让每个<p>标签的文本内容自动用换行分隔,同时移除所有标签结构:
from bs4 import BeautifulSoup with open('nosleep2.html', encoding='utf-8') as webpage: soup = BeautifulSoup(webpage, 'lxml') post = soup.find('div', class_='RichTextJSON-root') # 使用separator参数指定换行符分隔各段落 print(post.get_text(separator='\n'))
方法二:遍历所有
标签拼接文本
先定位所有<p>标签,提取每个标签的文本后用换行符拼接,适合需要额外处理单个段落的场景:
from bs4 import BeautifulSoup with open('nosleep2.html', encoding='utf-8') as webpage: soup = BeautifulSoup(webpage, 'lxml') post = soup.find('div', class_='RichTextJSON-root') # 获取所有p标签 p_paragraphs = post.find_all('p') # 拼接每个p标签的文本,用换行分隔 formatted_content = '\n'.join([p.get_text() for p in p_paragraphs]) print(formatted_content)
这两种方法都能输出你期望的排版效果,既移除了所有<p>标签及其属性,又让每个段落单独占一行。
内容的提问来源于stack exchange,提问作者Eliott
相关产品推荐
相关产品推荐

