You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何去除Python爬虫抓取文本中的大量空白?

解决爬虫结果大量空白的方案

你的问题出在直接调用table.get_text()时,会保留标签里所有的换行、空格;而string()方法只适用于单个文本节点的标签,table包含多层子标签,所以string()根本取不到内容,自然无效。下面是几种实用的解决方法:

方法1:用stripped_strings生成器

这个方法会自动过滤文本前后的空白和空字符串,直接输出干净内容:

from bs4 import BeautifulSoup
import requests

url = 'https://www.rucoyonline.com/characters/Something' 
response = requests.get(url)
print(response.status_code)

soup = BeautifulSoup(response.text, 'html.parser')

table = soup.find('table', class_ = 'character-table table table-bordered')
# 遍历stripped_strings输出每个清理后的文本块
for text in table.stripped_strings:
    print(text)

输出效果:

Character Information
Name
Something
Level
28
Last online
about 6 years ago
Born
September 03, 2016

方法2:给get_text()加参数

直接调用get_text()时,指定strip=True和separator='\n',可以一次性合并空白并按换行分隔:

# 替换原print语句
clean_content = table.get_text(strip=True, separator='\n')
print(clean_content)

代码更简洁,输出结果和方法1完全一致。

方法3:正则表达式替换(灵活定制)

如果需要更灵活的空白处理规则,比如把连续空白换成单个空格或换行,用正则表达式:

import re

raw_content = table.get_text()
# 把任意数量的空白(换行、空格、制表符等)替换成单个换行,再去掉首尾空白
clean_content = re.sub(r'\s+', '\n', raw_content).strip()
print(clean_content)

内容的提问来源于stack exchange,提问作者mr. one

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 19:45:39