You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式移除数据框stname列末尾的省略号?

解决数据框列末尾Unicode省略号移除问题

问题原因

你之前的方法失效,是因为数据里的“省略号”并非普通半角点(.,U+002E),而是Unicode水平省略号(…,U+2026),常规正则无法匹配这类特殊字符。

解决方案

以下是几种可行的处理方式,基于tidyverse的stringr包:

1. 直接移除所有Unicode省略号

针对数据中的连续…,直接匹配替换:

library(tidyverse)

df = structure(list(stname = c("Alabama……………………………………", 
"Alaska………………………………………", "American Samoa……………………………", 
"Arizona………………………………………", "Arkansas……………………………………", 
"California………………………………"), value = c(34305795, 
20236292, 103657, 267021650, 15045025, 3976908430)), row.names = c(NA, 
-6L), class = c("tbl_df", "tbl", "data.frame"))

# 替换所有连续的Unicode省略号
df_cleaned <- df %>%
  mutate(stname = str_remove_all(stname, "…+"))

2. 移除所有非字母数字/空格的Unicode字符

如果还存在其他特殊标点,可直接过滤掉所有非合法字符(支持Unicode字母、数字、空格):

df_cleaned <- df %>%
  mutate(stname = str_remove_all(stname, "[^\\p{L}\\p{N}\\s]"))

3. 仅移除末尾的省略号

如果需要保留字符串中间的合法标点,只清理末尾的连续省略号:

df_cleaned <- df %>%
  mutate(stname = str_replace(stname, "…+$", ""))

验证结果

执行后查看处理后的stname列:

print(df_cleaned$stname)
# 输出:
# [1] "Alabama"        "Alaska"         "American Samoa" "Arizona"        "Arkansas"       "California"

内容的提问来源于stack exchange,提问作者Mel G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 07:25:20