如何通过保留指定小数位缩小R Plotly HTML输出文件体积
问题描述
我正尝试缩小包含大量千级数据点Plotly图表的RMarkdown HTML报告体积、提升打开速度。由于R Plotly会把所有原始数据存入HTML,本以为通过保留小数位能减小体积,但发现即使对输入数据做了保留小数位处理,Plotly仍在HTML中存储大量小数位,文件体积没有变化。
测试案例1:原始数据生成的HTML
RawData <- data.frame(Date = seq(as.Date("2024/1/1"), by = "month", length.out = 12), PreciseValue = c(0.1516270, 0.3542629, 0.8339342, 0.5796813, 0.3933472, 0.2937137, 0.1779205, 0.4285533, 0.6841885, 0.3399411,0.99476560, 0.42941527)) RawData$RoundValue <- round(RawData$PreciseValue,2) fig <- plot_ly(RawData, type = 'scatter', mode = 'lines')%>% add_trace(x = ~Date, y = ~PreciseValue, name = 'PreciseValue') saveWidget(fig, "plotly_base.html", selfcontained = TRUE)
该HTML文件大小为3780kb,查看底层存储的y数据,小数位比原始数据更多:
"y":[0.15162704353000001,0.35426295622999998,0.83393426323999997,0.57968136341999998,0.39334726234,0.29371352347000002,0.17792423404999999,0.44352285533000002,0.68418423485000002,0.36623994110000002,0.99476432455999997,0.42941523452699998]
测试案例2:保留小数位数据生成的HTML
RawData$RoundValue <- round(RawData$PreciseValue,2) fig <- plot_ly(RawData, type = 'scatter', mode = 'lines')%>% add_trace(x = ~Date, y = ~RoundValue, name = 'RoundValue') saveWidget(fig, "plotly_round.html", selfcontained = TRUE)
该HTML文件大小同样为3780kb,底层存储的y数据仍存在大量冗余小数位:
"y":[0.14999999999999999,0.34999999999999998,0.82999999999999996,0.57999999999999996,0.39000000000000001,0.28999999999999998,0.17999999999999999,0.44,0.68000000000000005,0.37,0.98999999999999999,0.42999999999999999]
预期存储的y数据应为:
"y":[0.15, 0.35, 0.83, 0.58, 0.39, 0.29, 0.18, 0.44, 0.68, 0.37, 0.99, 0.43]
解决方案
问题根源在于浮点数的精度特性,以及Plotly默认的JSON序列化方式会保留完整的浮点数精度。要让Plotly在HTML中仅存储指定位数的小数,可以通过以下两种方法实现:
方法1:修改Plotly对象数据+控制JSON序列化精度
直接修改Plotly对象中的数据,将数值按指定小数位处理,同时在保存时指定JSON序列化的小数位数:
library(plotly) library(htmlwidgets) library(jsonlite) RawData <- data.frame(Date = seq(as.Date("2024/1/1"), by = "month", length.out = 12), PreciseValue = c(0.1516270, 0.3542629, 0.8339342, 0.5796813, 0.3933472, 0.2937137, 0.1779205, 0.4285533, 0.6841885, 0.3399411,0.99476560, 0.42941527)) RawData$RoundValue <- round(RawData$PreciseValue,2) # 生成图表 fig <- plot_ly(RawData, type = 'scatter', mode = 'lines')%>% add_trace(x = ~Date, y = ~RoundValue, name = 'RoundValue') # 修改data中的y值,用sprintf避免浮点数精度问题 fig$x$data[[1]]$y <- as.numeric(sprintf("%.2f", fig$x$data[[1]]$y)) # 保存时指定JSON序列化的小数位数 saveWidget( fig, "plotly_optimized.html", selfcontained = TRUE, preHook = function(widget) { widget$x <- toJSON(widget$x, digits = 2, auto_unbox = TRUE) widget } )
方法2:预处理数据+指定序列化参数
先将数据格式化为指定小数位的数值(通过sprintf规避浮点数精度问题),再生成图表,同时确保JSON序列化时使用指定精度:
library(plotly) library(htmlwidgets) library(jsonlite) RawData <- data.frame(Date = seq(as.Date("2024/1/1"), by = "month", length.out = 12), PreciseValue = c(0.1516270, 0.3542629, 0.8339342, 0.5796813, 0.3933472, 0.2937137, 0.1779205, 0.4285533, 0.6841885, 0.3399411,0.99476560, 0.42941527)) # 格式化数据为精确的2位小数 RawData$FormattedValue <- as.numeric(sprintf("%.2f", RawData$PreciseValue)) fig <- plot_ly(RawData, type = 'scatter', mode = 'lines')%>% add_trace(x = ~Date, y = ~FormattedValue, name = 'FormattedValue') saveWidget( fig, "plotly_optimized2.html", selfcontained = TRUE, preHook = function(widget) { widget$x <- toJSON(widget$x, digits = 2, auto_unbox = TRUE) } )
效果验证
使用上述方法后,HTML文件中的y数据会变为预期的精简格式,文件体积会显著减小,同时图表的视觉效果不受影响。
内容的提问来源于stack exchange,提问作者Anthony
相关产品推荐
相关产品推荐

