You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言数据框的字符列中提取response rate对应的百分比数值

R语言提取seller_details列中response rate对应数值的实现方法

核心逻辑:通过正则匹配定位response rate后、%前的数字片段,提取后转换为数值类型即可实现需求。


方法1:使用stringr包(tidyverse生态,代码更简洁)

适合日常使用tidyverse工具链的场景,正则断言直接提取目标片段:

# 加载依赖包
library(stringr)
library(dplyr)

# 导入测试数据
df <- structure(list(seller_details = c("8ysl9a1301 Active 3 hours ago chat now view shop Ratings10products996 response rate28% response", 
"showcasemywardore Active 3 hours ago chat now view shop Ratings773products5k response rate70% response", 
"zanzea.os Active 37 minutes ago chat now view shop Ratings290.5kproducts6.6k response rate93% response", 
"airspacemy.os Active 14 minutes ago chat now view shop Ratings1.2kproducts2k response rate70% response", 
"zanzea.os Active 37 minutes ago chat now view shop Ratings290.5kproducts6.6k response rate93% response"
)), class = "data.frame", row.names = c(NA, -5L))

# 提取数值并存入新列
df <- df %>%
  mutate(response_rate = as.integer(str_extract(seller_details, pattern = "(?<=response rate)\\d+(?=%)")))

正则说明:(?<=response rate)是正向后向断言,匹配前缀为response rate的位置;\\d+匹配1个及以上数字;(?=%)是正向前向断言,匹配后缀为%的位置,三者组合直接定位到目标数字。


方法2:Base R实现(无需额外安装包)

适合不想安装第三方依赖的场景,用字符串替换逻辑实现:

# 沿用上方的测试数据df
df$response_rate <- as.integer(sub(".*response rate(\\d+)%.*", "\\1", df$seller_details))

逻辑说明:sub函数会匹配整段字符串,捕获response rate和%之间的数字组,再用捕获到的第一组内容替换整个字符串,得到纯数字字符串后转换为整数即可。

两种方法输出的response_rate列结果均为:28 70 93 70 93,完全符合预期。


内容的提问来源于stack exchange,提问作者Mohd Syazwan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 11:57:02