You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求R语言中高效转换中数混合字符串为英数混合字符串的方案

Absolutely! R has got you covered for converting these mixed Chinese-numeric strings to English-numeric ones efficiently—pattern matching and replacement are your best friends here. Let’s break this down into your specific use cases:

1. Converting Date Columns ("--月--日" → "-- month -- day")

Both base R and tidyverse tools work great here, with no need for complex loops.

Base R (no extra packages required)

Use gsub() with capturing groups to grab numeric values and reformat them with English terms:

# Assume your data frame is named df, and the date column is called date_col
df$date_col <- gsub("(\\d+)月(\\d+)日", "\\1 month \\2 day", df$date_col)

Here, (\\d+) captures one or more digits, and \\1/\\2 refers back to those captured numbers in the replacement string—super flexible for this kind of pattern swap.

Tidyverse (cleaner syntax with stringr)

If you’re already using the tidyverse ecosystem, str_replace_all() makes this even more intuitive:

library(stringr)

df$date_col <- str_replace_all(df$date_col, "(\\d+)月(\\d+)日", "\\1 month \\2 day")

2. Converting Deviation Date Columns

For "提前--天" → "advance -- days" and "推迟--天" → "postpone -- days", you can handle both cases efficiently:

Base R

Chain two gsub() calls to cover both patterns sequentially:

# Assume deviation column is named deviation_col
df$deviation_col <- gsub("提前(\\d+)天", "advance \\1 days", df$deviation_col)
df$deviation_col <- gsub("推迟(\\d+)天", "postpone \\1 days", df$deviation_col)

Tidyverse (one-step multi-pattern replacement)

str_replace_all() accepts a named vector to handle multiple replacement rules in a single pass:

library(tidyverse)

df <- df %>%
  mutate(deviation_col = str_replace_all(deviation_col, 
                                         c("提前(\\d+)天" = "advance \\1 days", 
                                           "推迟(\\d+)天" = "postpone \\1 days")))

Quick Note

All these functions are vectorized, meaning they process your entire column in one go—no slow loops, even for large datasets. If you run into similar mixed-string patterns later, just extend the regex: \\d+ will always capture numeric values, and you can tweak the replacement terms to match new cases.

内容的提问来源于stack exchange,提问作者Daniel Luon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:47:42