网页抓取后如何去除字符串中的\r\n转义码?Python数据清洗问题咨询
解决网页抓取结果中的\r\n转义字符问题
嘿,我来帮你搞定这个问题!你遇到的错误是因为replace和strip都是字符串方法,不能直接用在列表上——你的data是嵌套列表结构,得逐个处理里面的字符串元素才行。
两种解决方案:
1. 抓取数据时直接清理(推荐)
在收集每个单元格文本的时候就做清理,这样后续不用再额外处理。修改你代码中遍历cell的部分:
import pandas as pd import numpy as np import matplotlib.pyplot as plt %matplotlib inline from urllib.request import urlopen from bs4 import BeautifulSoup import re url="https://www.hubertiming.com/results/2018MLK" # OPEN LINK html=urlopen(url) # 注意你原来写的是大写URL,这里要和变量名保持一致 soup=BeautifulSoup(html,"lxml") title = soup.title print(title) print(title.text) links = soup.find_all('a',href=True) for link in links: print(link['href']) data =[] allrows=soup.find_all("tr") for row in allrows: row_list = row.find_all("td") dataRow=[] for cell in row_list: # 关键修改:用strip()清理字符串首尾的空白、换行、回车 cleaned_text = cell.text.strip() dataRow.append(cleaned_text) data.append(dataRow) data=data[4:] print(data[-2:])
2. 对已有的data列表事后清理
如果已经拿到了带冗余字符的data,可以用嵌套列表推导式批量处理:
# 遍历每个子列表,再遍历每个字符串元素做清理 cleaned_data = [[item.strip() for item in row] for row in data] print(cleaned_data[-2:])
为什么这样有效?
strip()方法会自动移除字符串首尾的所有空白类字符,包括:
\r(回车符)\n(换行符)- 空格、制表符
\t等
正好完美解决你数据里的冗余字符问题,同时还能清理掉名字前后的多余空格(比如\r\n\r\n LEESHA POSEY\r\n\r\n 会变成LEESHA POSEY)。
处理后你会得到干净的结果:
[['190', '2087', 'LEESHA POSEY', 'F', '43', 'PORTLAND', 'OR', '1:33:53', '30:17', '112 of 113', 'F 40-54', '36 of 37', '0:00', '1:33:53'], ['191', '1216', 'ZULMA OCHOA', 'F', '40', 'GRESHAM', 'OR', '1:43:27', '33:22', '113 of 113', 'F 40-54', '37 of 37', '0:00', '1:43:27']]
内容的提问来源于stack exchange,提问作者andrila
相关产品推荐
相关产品推荐

