You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用R干净解析HTML文件中的表格,解决结果含=\r\n的问题

R解析HTML表格多余=\r\n字符解决方案

你遇到的=\r\n是*可打印引用编码(Quoted-Printable,QP编码)*的软换行标记,通常出现在特定传输场景下的HTML文件中,属于编码层面的特殊标记,可通过以下两种方案解决:


方案1:预处理HTML原始文本再解析(推荐)

在读取HTML文件时先替换所有QP软换行,再解析表格,从根源避免脏数据进入解析结果:

library(tidyverse)
library(xml2)
library(rvest)
library(janitor)

# 读取原始HTML文本,批量删除所有QP软换行
html_raw <- readLines("my_html_file.html", encoding = "UTF-8", warn = FALSE) %>%
  # 正则匹配=加Windows换行(\r\n)或Unix换行(\n),直接删除
  str_replace_all("=\\r?\\n", "")

# 将处理后的文本转换为HTML对象
html <- read_html(paste(html_raw, collapse = "\n"),
                        encoding = "UTF-8",
                        options = "RECOVER",
                        warn = FALSE, verbose = TRUE)
html_res <- rvest::html_table(html, trim = FALSE)

df <- html_res[[1]] %>%
  janitor::clean_names() %>%
  dplyr::select(run_id, s_id, s_number, library, s_name, tissue)

方案2:解析后批量清理字符列

如果已经完成表格解析,也可以直接对所有字符列批量过滤脏字符:

df_clean <- df %>%
  janitor::clean_names() %>%
  dplyr::select(run_id, s_id, s_number, library, s_name, tissue) %>%
  # 遍历所有字符列,删除所有=\r\n标记
  dplyr::mutate(dplyr::across(where(is.character), ~stringr::str_remove_all(.x, "=\\r?\\n")))

扩展:完整QP编码解码

如果还存在其他QP编码残留(比如=3D代表等号、=20代表空格等),可以直接使用QP解码函数处理原始文本,一步清理所有编码标记:

# 需先安装qdap包:install.packages("qdap")
library(qdap)

html_raw <- readLines("my_html_file.html", encoding = "UTF-8", warn = FALSE) %>%
  paste(collapse = "\n") %>%
  qp_decode()

html <- read_html(html_raw, encoding = "UTF-8")

内容的提问来源于stack exchange,提问作者littleworth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 07:12:01