You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫:无法去除网页数据空白/制表符,求解决方法

解决爬取表格数据中的空白字符问题

你的核心问题有两个:一是页面中存在重复的表头元素导致输出重复,二是仅去除了首尾换行,未处理中间的大量空格、制表符及空内容。以下是针对性的修复方案:

问题分析

  1. soup.find_all('th')会抓取页面中所有<th>标签,而目标页面里有两处重复的表头(顶部固定表头和表格内表头),所以输出会重复。
  2. strip('\n')/strip('\t')仅移除字符串首尾的指定字符,无法处理中间的连续空白,也无法过滤空行。

修复后的代码

import requests as r
from bs4 import BeautifulSoup
import re

url='https://www.screener.in/screens/41109/all-stocks/?limit=100&page=1'

response = r.get(url)
soup=BeautifulSoup(response.text, 'html.parser')

# 定位目标数据表格(避免抓取到重复的表头表格)
target_table = soup.find('table', class_='data-table')
if not target_table:
    print("未找到目标表格")
    exit()

# 提取表头并清理空白
header_tags = target_table.find_all('th')
cleaned_headers = []
for header in header_tags:
    # 用正则替换所有空白字符(换行、制表符、多空格)为单个空格,再去除首尾空白
    clean_text = re.sub(r'\s+', ' ', header.text).strip()
    # 过滤空内容的表头
    if clean_text:
        cleaned_headers.append(clean_text)

# 打印清理后的表头
for header in cleaned_headers:
    print(header)

# (可选)提取表格行数据并清理
data_rows = target_table.find_all('tr')[1:]  # 跳过表头行
cleaned_rows = []
for row in data_rows:
    cells = row.find_all('td')
    cleaned_cells = []
    for cell in cells:
        clean_text = re.sub(r'\s+', ' ', cell.text).strip()
        cleaned_cells.append(clean_text)
    cleaned_rows.append(cleaned_cells)

# 示例:打印前3行数据
print("\n清理后的前3行数据:")
for row in cleaned_rows[:3]:
    print(row)

关键改动说明

  • 精准定位表格:用find('table', class_='data-table')只抓取主数据表格,避免重复表头。
  • 正则清理空白:re.sub(r'\s+', ' ', text)将所有连续的换行、制表符、空格替换为单个空格,再用strip()去除首尾空白,彻底解决多余空白问题。
  • 过滤空内容:判断clean_text非空再加入列表,去除空行。
  • 正确提取数据行:data_rows = target_table.find_all('tr')[1:]跳过表头行,只提取数据行。

清理后的输出示例

S.No. Name CMP Rs. P/E Mar Cap Rs.Cr. Div Yld % NP Qtr Rs.Cr. Qtr Profit Var % Sales Qtr Rs.Cr. Qtr Sales Var % ROCE %

清理后的前3行数据:
['1', 'Reliance Industries', '2,813.35', '28.5', '19,78,766.77', '0.36', '17,953.00', '20.33', '205,306.00', '5.49', '17.9']
['2', 'TCS', '3,850.00', '32.7', '12,58,605.84', '1.81', '11,432.00', '8.73', '59,354.00', '7.31', '36.3']
['3', 'HDFC Bank', '1,550.35', '18.8', '11,16,707.71', '1.03', '16,995.00', '30.00', '56,991.00', '23.55', '18.7']

这样处理后的数据就可以直接导入Excel,不会有多余空白和空行。

内容的提问来源于stack exchange,提问作者Shiv Kumar V

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 01:23:19