如何移除R数据框col1列条目的末尾下划线
移除R数据框指定列末尾的下划线(仅匹配末尾字符)
需求说明
仅当数据框col1列的字符串末尾字符为下划线_时,移除该下划线;字符串中间的下划线保留不变。
示例数据
data1 <- c("foo_bar_","bar_foo","apple_","apple__beer_") df <- data.frame("col1"=data1,"col2"=1:4)
当前输出:
col1 col2 1 foo_bar_ 1 2 bar_foo 2 3 apple_ 3 4 apple__beer_ 4
期望输出:
col1 col2 1 foo_bar 1 2 bar_foo 2 3 apple 3 4 apple__beer 4
解决方案
方法1:基础R原生函数gsub
利用正则表达式匹配末尾的下划线,替换为空字符串:
df$col1 <- gsub("_$", "", df$col1)
- 正则说明:
_$中的$表示匹配字符串的末尾位置,因此只会替换最后一个字符是下划线的情况,不会影响字符串中间的下划线。
方法2:使用stringr包(语法更直观)
如果习惯tidyverse风格,可以用stringr包的str_remove函数:
# 未安装包时先执行:install.packages("stringr") library(stringr) df$col1 <- str_remove(df$col1, "_$")
- 逻辑和基础R方法一致,
str_remove会精准移除匹配到的末尾下划线。
执行任意一种方法后,查看df即可得到期望的结果。
内容的提问来源于stack exchange,提问作者Geli
相关产品推荐
相关产品推荐

