You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium+BeautifulSoup爬虫 如何统一两站点抓取文本的格式

问题原因

两个网站的DOM排版逻辑不同,第二个站点大量依赖div、p、标题、列表项等块级元素的默认换行实现排版,你当前代码直接使用getText(separator=u' ')提取文本时,只会统一用空格分隔所有节点内容,丢失了块级元素原本的换行分隔属性,最终导致内容全部挤在一起。

解决方法

先遍历DOM给所有块级元素的末尾追加换行符,再提取文本,最后统一处理多余的空白字符即可,修改后的代码如下:

from bs4 import BeautifulSoup
from selenium import webdriver
import urllib.parse
from selenium.common.exceptions import WebDriverException
from selenium.webdriver.chrome.service import Service
import os
import re

service = Service("/home/ubuntu/selenium_drivers/chromedriver")

options = webdriver.ChromeOptions()
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.131 Safari/537.3")
options.add_argument("--headless")
options.add_argument('--ignore-certificate-errors')
options.add_argument("--enable-javascript")
options.add_argument('--incognito')

# 两个地址都可以直接替换测试
URL = "https://msorchestra.com/event/41st-annual-pepsi-pops-a-blast-in-the-park-3/"
# URL = "https://www.ncco.org/2021-season/set-ii-call-of-destiny"

try:
    driver = webdriver.Chrome(service = service, options = options)
    driver.get(URL)
    driver.implicitly_wait(2)
    html_content = driver.page_source
    driver.quit()
except WebDriverException:
    driver.quit()

soup = BeautifulSoup(html_content, 'html.parser')

for h in soup.find_all('header'):
    try:
        h.extract()
    except:
        pass
for f in soup.find_all('footer'):
    try:
        f.extract()
    except:
        pass

# 新增:给所有块级元素末尾追加换行符,保留排版逻辑
block_tags = ['p', 'div', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6', 'li', 'tr', 'br', 'section', 'article']
for tag in block_tags:
    for element in soup.find_all(tag):
        element.append('\n')

# 提取文本后处理多余空白
text = soup.get_text()
# 合并连续的空格,保留换行
text = re.sub(r' +', ' ', text)
# 合并连续超过2个的换行,避免大量空行
text = re.sub(r'\n\s*\n', '\n\n', text)
# 去掉每行首尾的空白
text = '\n'.join([line.strip() for line in text.split('\n')])

print(text)

调整说明

  • 新增了常见块级标签的遍历逻辑,给每个块级元素末尾主动插入换行,保留页面原本的排版层级
  • 新增了正则处理逻辑,既保留了必要的换行、空格,又避免了冗余空白导致的排版混乱
  • 对两个站点的适配性一致,输出的文本可读性不会再出现明显差异

内容的提问来源于stack exchange,提问作者imhans33

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 18:36:03