使用Selenium+BeautifulSoup爬虫 如何统一两站点抓取文本的格式
问题原因
两个网站的DOM排版逻辑不同,第二个站点大量依赖div、p、标题、列表项等块级元素的默认换行实现排版,你当前代码直接使用getText(separator=u' ')提取文本时,只会统一用空格分隔所有节点内容,丢失了块级元素原本的换行分隔属性,最终导致内容全部挤在一起。
解决方法
先遍历DOM给所有块级元素的末尾追加换行符,再提取文本,最后统一处理多余的空白字符即可,修改后的代码如下:
from bs4 import BeautifulSoup from selenium import webdriver import urllib.parse from selenium.common.exceptions import WebDriverException from selenium.webdriver.chrome.service import Service import os import re service = Service("/home/ubuntu/selenium_drivers/chromedriver") options = webdriver.ChromeOptions() options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.131 Safari/537.3") options.add_argument("--headless") options.add_argument('--ignore-certificate-errors') options.add_argument("--enable-javascript") options.add_argument('--incognito') # 两个地址都可以直接替换测试 URL = "https://msorchestra.com/event/41st-annual-pepsi-pops-a-blast-in-the-park-3/" # URL = "https://www.ncco.org/2021-season/set-ii-call-of-destiny" try: driver = webdriver.Chrome(service = service, options = options) driver.get(URL) driver.implicitly_wait(2) html_content = driver.page_source driver.quit() except WebDriverException: driver.quit() soup = BeautifulSoup(html_content, 'html.parser') for h in soup.find_all('header'): try: h.extract() except: pass for f in soup.find_all('footer'): try: f.extract() except: pass # 新增:给所有块级元素末尾追加换行符,保留排版逻辑 block_tags = ['p', 'div', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6', 'li', 'tr', 'br', 'section', 'article'] for tag in block_tags: for element in soup.find_all(tag): element.append('\n') # 提取文本后处理多余空白 text = soup.get_text() # 合并连续的空格,保留换行 text = re.sub(r' +', ' ', text) # 合并连续超过2个的换行,避免大量空行 text = re.sub(r'\n\s*\n', '\n\n', text) # 去掉每行首尾的空白 text = '\n'.join([line.strip() for line in text.split('\n')]) print(text)
调整说明
- 新增了常见块级标签的遍历逻辑,给每个块级元素末尾主动插入换行,保留页面原本的排版层级
- 新增了正则处理逻辑,既保留了必要的换行、空格,又避免了冗余空白导致的排版混乱
- 对两个站点的适配性一致,输出的文本可读性不会再出现明显差异
内容的提问来源于stack exchange,提问作者imhans33
相关产品推荐
相关产品推荐

