You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取后如何去除字符串中的\r\n转义码?Python数据清洗问题咨询

解决网页抓取结果中的\r\n转义字符问题

嘿,我来帮你搞定这个问题!你遇到的错误是因为replace和strip都是字符串方法,不能直接用在列表上——你的data是嵌套列表结构,得逐个处理里面的字符串元素才行。

两种解决方案:

1. 抓取数据时直接清理(推荐)

在收集每个单元格文本的时候就做清理,这样后续不用再额外处理。修改你代码中遍历cell的部分:

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
%matplotlib inline
from urllib.request import urlopen
from bs4 import BeautifulSoup
import re

url="https://www.hubertiming.com/results/2018MLK"
# OPEN LINK
html=urlopen(url)  # 注意你原来写的是大写URL,这里要和变量名保持一致
soup=BeautifulSoup(html,"lxml")
title = soup.title
print(title)
print(title.text)

links = soup.find_all('a',href=True)
for link in links:
    print(link['href'])

data =[]
allrows=soup.find_all("tr")
for row in allrows:
    row_list = row.find_all("td")
    dataRow=[]
    for cell in row_list:
        # 关键修改:用strip()清理字符串首尾的空白、换行、回车
        cleaned_text = cell.text.strip()
        dataRow.append(cleaned_text)
    data.append(dataRow)

data=data[4:]
print(data[-2:])

2. 对已有的data列表事后清理

如果已经拿到了带冗余字符的data,可以用嵌套列表推导式批量处理:

# 遍历每个子列表,再遍历每个字符串元素做清理
cleaned_data = [[item.strip() for item in row] for row in data]
print(cleaned_data[-2:])

为什么这样有效?

strip()方法会自动移除字符串首尾的所有空白类字符,包括:

  • \r(回车符)
  • \n(换行符)
  • 空格、制表符\t等

正好完美解决你数据里的冗余字符问题,同时还能清理掉名字前后的多余空格(比如\r\n\r\n LEESHA POSEY\r\n\r\n 会变成LEESHA POSEY)。

处理后你会得到干净的结果:

[['190', '2087', 'LEESHA POSEY', 'F', '43', 'PORTLAND', 'OR', '1:33:53', '30:17', '112 of 113', 'F 40-54', '36 of 37', '0:00', '1:33:53'], ['191', '1216', 'ZULMA OCHOA', 'F', '40', 'GRESHAM', 'OR', '1:43:27', '33:22', '113 of 113', 'F 40-54', '37 of 37', '0:00', '1:43:27']]

内容的提问来源于stack exchange,提问作者andrila

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 15:47:43