如何去除BeautifulSoup爬取数据中多余的空格、换行符
问题原因
- 你调用
BeautifulSoup(html, 'lxml').text之后已经不存在HTML标签了,后续用正则<.*?>替换标签属于无效操作 - 你提前把所有换行符
\n、制表符\t都替换成了空格,导致文本里所有的排版间隔都变成了连续空格,strip()仅能去除整个字符串首尾的空格,无法处理中间的连续空格,也没法处理你后续拆分出来的每行首尾空格 - 要处理行首空格的前提是先把文本按行拆分,逐行处理,而不是先把所有换行都换成空格抹掉行结构
解决方案
场景1:需要得到无换行、无多余空格的整段字符串
直接用正则替换所有连续空白符为单个空格,再整体trim即可:
from bs4 import BeautifulSoup import requests import re URL = "https://www.flagstaffsymphony.org/event/a-flag-on-fourth/" html_content = requests.get(URL).text # 提取纯文本 cleantext = BeautifulSoup(html_content, "lxml").text # 先过滤非ASCII字符 text = re.sub(r'[^\x00-\x7F]+', ' ', cleantext) # 把所有连续空白符(空格、换行、制表等)替换为单个空格,再去除首尾空格 text = re.sub(r'\s+', ' ', text).strip() with open('read.txt', 'w', encoding='utf-8') as file: file.write(text)
场景2:需要保留换行结构,每行无首尾空格、无空行
逐行处理后过滤空行再拼接:
from bs4 import BeautifulSoup import requests import re URL = "https://www.flagstaffsymphony.org/event/a-flag-on-fourth/" html_content = requests.get(URL).text cleantext = BeautifulSoup(html_content, "lxml").text # 先替换非ASCII字符为空格 cleantext = re.sub(r'[^\x00-\x7F]+', ' ', cleantext) # 按行拆分,逐行去除首尾空格,过滤空行 lines = [line.strip() for line in cleantext.splitlines() if line.strip()] # 拼接成带换行的文本 text = '\n'.join(lines) with open('read.txt', 'w', encoding='utf-8') as file: file.write(text)
单独处理你给出的样例数据
直接逐行strip()即可:
sample = """Subscription Tickets All Events This event has passed. America the Beautiful: A Virtual Patriotic Salute July 4, 2020 Violin Virtuoso Beethoven Virtual 5k """ processed = '\n'.join([line.strip() for line in sample.splitlines()]) print(processed)
处理后即可得到无行首行尾多余空格的内容。
内容的提问来源于stack exchange,提问作者Aniiya0978
相关产品推荐
相关产品推荐

