You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R的arrow包读取带特殊格式数值的大型TXT文件?

解决Arrow包读取特殊格式数值的大文件问题

你的问题核心是Arrow的默认解析逻辑不支持右侧负号、点作为千位分隔符、逗号作为小数点的特殊数值格式,所以需要先读取为字符串,再手动转换为数值类型,分两种场景处理:

一、内存充足时的处理

先将所有列按字符类型读取,再针对数值列做格式转换:

  1. 按字符类型读取文件
library(arrow)
library(dplyr)

# 强制指定数值列为字符串类型,避免Arrow自动解析失败
df <- read_delim_arrow(
  file = "./Dati/esempio.txt",
  delim = "|",
  col_types = schema(
    text01 = string(),
    number01 = string(),
    number02 = string(),
    number03 = string()
  )
)
  1. 定义转换函数并处理数值列
# 自定义函数:处理特殊格式数值
convert_special_numeric <- function(x) {
  # 把右侧负号移到开头
  x <- ifelse(grepl("-$$", x), paste0("-", substr(x, 1, nchar(x)-1)), x)
  # 移除千位分隔符的点,将逗号替换为标准小数点
  x <- gsub("\\.", "", x)
  x <- gsub(",", ".", x)
  # 转换为数值类型
  as.numeric(x)
}

# 批量处理所有以number开头的列
df_processed <- df %>%
  mutate(across(starts_with("number"), convert_special_numeric))

二、内存不足时的Dataset处理

如果文件过大无法加载到内存,用open_delim_dataset创建数据集,结合Arrow支持的dplyr语法做流式处理:

  1. 创建数据集对象
ds <- open_delim_dataset(
  "./Dati/esempio.txt",
  delim = "|",
  col_types = schema(
    text01 = string(),
    number01 = string(),
    number02 = string(),
    number03 = string()
  )
)
  1. 流式转换并导出(避免加载到内存)
# 使用Arrow兼容的字符串函数处理格式
ds_processed <- ds %>%
  mutate(
    across(starts_with("number"), ~ {
      # 处理右侧负号
      case_when(
        str_ends(., "-") ~ paste0("-", str_sub(., 1, str_length(.) - 1)),
        TRUE ~ .
      ) %>%
        str_replace_all("\\.", "") %>%  # 移除千位分隔符
        str_replace(",", ".") %>%      # 替换小数点格式
        as.numeric()                   # 转换为数值
    })
  )

# 将处理后的数据集导出到磁盘(可选择parquet等高效格式)
write_dataset(ds_processed, path = "./processed_data", format = "parquet")

关键说明

  • Arrow默认的数值解析只支持标准格式(负号在左、千位分隔符可选但需指定、小数点为点或指定符号),无法处理右侧负号的情况,所以必须手动做字符串转换。
  • 使用Dataset方式时,所有转换操作都是在磁盘上流式执行的,不会一次性加载全部数据,适合超大型文件。

内容的提问来源于stack exchange,提问作者Dmozzanica

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:54:57