You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup将爬取文本中的换行符替换为句号加空格?

解决方案:替换换行符为句号加空格

我来帮你搞定这个需求!你现在的问题是原页面的div.description内容里存在换行,用get_text(separator=" ")会把换行直接转成空格,导致文本生硬连在一起,你需要把这些原换行的位置替换成. (句号加空格)来让语句更通顺。

问题根源

你原来的代码用了get_text(separator=" "),它会把div里所有的换行、制表符等空白字符统一替换成单个空格,这样就没法区分原文本里的换行分隔位置了。要实现目标,得先保留原始的换行符,把它们替换成. 之后再处理多余的空格。

修改后的代码

这里提供两种可行方案,你可以根据偏好选择:

方案1:正则表达式(推荐,处理更灵活)

import sys
import urllib2
from bs4 import BeautifulSoup
import re  # 导入正则模块

quote_page = sys.argv[1]
page = urllib2.urlopen(quote_page)
soup = BeautifulSoup(page, 'html.parser')
description_box = soup.find('div', {'class':'description'})

# 获取保留原始换行的文本
raw_text = description_box.get_text().strip()
# 把所有换行(包括前后的空格)替换成 ". "
description = re.sub(r'\s*\n\s*', '. ', raw_text)
# 合并多余的连续空格为单个空格
description = re.sub(r'\s+', ' ', description)

print(description)

方案2:纯字符串方法(无需正则)

import sys
import urllib2
from bs4 import BeautifulSoup

quote_page = sys.argv[1]
page = urllib2.urlopen(quote_page)
soup = BeautifulSoup(page, 'html.parser')
description_box = soup.find('div', {'class':'description'})

# 获取保留原始换行的文本
raw_text = description_box.get_text().strip()
# 替换换行符为 ". "
description = raw_text.replace('\n', '. ')
# 循环替换多个连续空格为单个空格
while '  ' in description:
    description = description.replace('  ', ' ')

print(description)

效果说明

运行修改后的脚本后,原文本中换行的位置会被精准替换成. ,最终输出就是你想要的:

Planet Nine was initially proposed to explain the clustering of orbits. Of Planet Nine's other effects, one was unexpected, the perpendicular orbits, and the other two were found after further analysis. Although other mechanisms have been offered for many of these peculiarities, the gravitational influence of Planet Nine is the only one that explains all four.

内容的提问来源于stack exchange,提问作者mumer91

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:01:59