You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将网页抓取数据转换为data.frame?如何清除\t\n\t\t\t空白字符?

解决网页抓取后空白字符清理的问题

我来帮你搞定这个烦人的空白字符问题!你遇到的\t、\n这类杂乱空白,其实用R里的工具很容易清理,给你两种实用方案:

方案一:用stringr包的str_squish()(推荐,最省心)

这个函数是tidyverse生态里的工具,能一键搞定所有多余空白:

  • 去除字符串首尾的所有空白(包括\t、\n、空格)
  • 把中间连续的任意空白字符(不管是哪种)替换成单个空格

完整代码示例:

# 加载所需包
library(rvest)
library(stringr)

# 定义目标URL
URL <- "https://rotogrinders.com/pages/pga-course-history-743469"

# 读取网页并抓取清理文本
tablescrape_html <- read_html(URL)
cleaned_text <- tablescrape_html %>% 
  html_nodes("tr") %>% 
  html_text() %>% 
  str_squish()  # 核心清理步骤

# 查看清理后的结果
head(cleaned_text)

方案二:用基础R函数(无需额外装包)

如果你不想安装新包,用基础R的trimws()和gsub()组合也能实现:

library(rvest)

URL <- "https://rotogrinders.com/pages/pga-course-history-743469"
tablescrape_html <- read_html(URL)

cleaned_text <- tablescrape_html %>% 
  html_nodes("tr") %>% 
  html_text() %>% 
  trimws() %>%  # 先去除首尾空白
  gsub("\\s+", " ", .)  # 把中间任意数量的空白换成单个空格

head(cleaned_text)

这里的\\s+正则表达式会匹配一个或多个任意空白字符(包括\t、\n、空格等),所以能一次性清理掉所有杂乱的空白格式。

内容的提问来源于stack exchange,提问作者user3066155

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:50:23