如何用BeautifulSoup将解析内容保存为无<p>标签的纯文本文件?
解决BeautifulSoup保存纯文本时保留HTML标签的问题
原脚本存在的问题
- 第一个脚本中,循环内打印
elem.text能正确输出纯文本,但最后执行file.write(str(elem))时,是将最后一个标签对象转为字符串写入,因此会保留<p>标签,且仅保存了最后一个元素的内容。 - 第二个修改脚本中,
data变量未定义,直接调用data.get_text()会报错;此外,将elem赋值为文本后,再调用elem.text会触发AttributeError,因为字符串类型没有text属性。
正确实现方案
方案一:直接提取目标div下的所有纯文本
这种方法会自动忽略div内所有HTML标签,提取全部纯文本内容:
from bs4 import BeautifulSoup # 读取HTML文件内容 with open('var/www/html/audi/index.html', 'r') as f: feed = f.read() soup = BeautifulSoup(feed, 'html.parser') # 定位目标div元素 target_div = soup.find("div", {"itemprop":"text"}) # 提取所有纯文本,用换行分隔段落,去除多余空白 pure_text = target_div.get_text(separator='\n', strip=True) # 将纯文本写入txt文件 with open("/var/www/html/audi/save.txt", "w") as file: file.write(pure_text)
方案二:遍历提取所有
标签的文本(更灵活)
如果只需提取div内<p>标签的内容,可逐个遍历并拼接:
from bs4 import BeautifulSoup with open('var/www/html/audi/index.html', 'r') as f: feed = f.read() soup = BeautifulSoup(feed, 'html.parser') target_div = soup.find("div", {"itemprop":"text"}) # 收集所有p标签的文本内容 text_content = [] for p_tag in target_div.find_all("p"): line = p_tag.get_text(strip=True) if line: # 跳过空内容的行 text_content.append(line) # 将列表内容转为字符串,用换行分隔 pure_text = '\n'.join(text_content) # 保存到文件 with open("/var/www/html/audi/save.txt", "w") as file: file.write(pure_text)
内容的提问来源于stack exchange,提问作者CodeBYa
相关产品推荐
相关产品推荐

