You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup爬虫时如何去除列表中的 冗余字符

爬取网页表格时多余换行、空白字符的解决方案

这类脏数据是BeautifulSoup提取文本时默认保留HTML源码里的排版格式字符导致的,从爬取提取环节或者已生成的DataFrame环节都可以快速清洗,不需要复杂处理。

源头清洗方案(推荐,爬取阶段直接生成干净数据)

  • 提取单元格文本时不要直接调用.text后就存储,增加两步字符串处理:
    1. 用正则把所有\r、\n、\t这类排版控制字符替换为单个空格
    2. 调用字符串自带的.strip()方法去掉首尾所有多余空白
  • 核心实现代码:
import urllib.request
from bs4 import BeautifulSoup
import pandas as pd
import re

# 读取目标页面
url = "https://www.hubertiming.com/results/2018MLK"
with urllib.request.urlopen(url) as response:
    page_content = response.read()
soup = BeautifulSoup(page_content, "html.parser")

table_rows = soup.find_all("tr")
cleaned_dataset = []
for tr in table_rows:
    # 同时提取表头th和数据单元格td
    cells = tr.find_all(["th", "td"])
    row_content = []
    for cell in cells:
        raw_txt = cell.text
        # 替换所有连续换行、回车、制表符为单个空格
        mid_txt = re.sub(r"[\r\n\t]+", " ", raw_txt)
        # 去除首尾多余空白
        final_txt = mid_txt.strip()
        row_content.append(final_txt)
    # 跳过无内容的空行
    if len(row_content) > 0:
        cleaned_dataset.append(row_content)

# 组装DataFrame,第一行为表头
df = pd.DataFrame(cleaned_dataset[1:], columns=cleaned_dataset[0])

运行后生成的DataFrame中,姓名、组内排名等所有字段都不会残留\r\n\r\n和多余缩进空格。

存量数据清洗方案(适用于已生成脏DataFrame的场景)

如果你已经完成爬取、生成了带脏数据的DataFrame,不需要重新爬取,直接批量处理所有字符串列即可:

# 遍历所有列,仅对字符串类型列做清洗
for column in df.columns:
    if pd.api.types.is_string_dtype(df[column]):
        df[column] = df[column].str.replace(r"[\r\n\t]+", " ", regex=True).str.strip()

注意:清洗时不要用无差别匹配所有空白的正则,避免删掉姓名、地址等字段中正常存在的词间空格,仅替换排版产生的控制类空白字符即可。

内容的提问来源于stack exchange,提问作者MalharK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 04:12:22