You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式从路径字符串提取Country、League、Year字段?

基于正则表达式从文件路径提取信息的R实现

原始数据集代码

path<-c("C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/argentina-primera-division-matches-2022-to-2022-stats.csv",
"C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/argentina-primera-division-matches-2022-to-2022-stats.csv",
"C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/france-ligue-2-matches-2021-to-2022-stats.csv",
"C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/germany-2-bundesliga-matches-2021-to-2022-stats.csv")

mydata<-data.frame(path=path)

需求

从path变量衍生三个新变量:

  • Country:提取DONNEES/后到第一个-之间的国家名称
  • League:提取DONNEES/后国家名称之后到-matches之前的联赛名称
  • Year:提取-to之前的年份信息

期望结果

path   Country           League Year
1 C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/argentina-primera-division-matches-2022-to-2022-stats.csv argentina primera-division 2022
2 C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/argentina-primera-division-matches-2022-to-2022-stats.csv argentina primera-division  202
3             C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/france-ligue-2-matches-2021-to-2022-stats.csv    france          ligue-2 2021
4       C:/Users/SEYDOU GORO/Dropbox/PC (2)/Documents/BETCLIC/DONNEES/germany-2-bundesliga-matches-2021-to-2022-stats.csv   germany     2-bundesliga 2021

正则表达式实现方法

方法1:Base R 原生函数 sub

利用sub的正则替换功能,捕获目标内容并提取:

# 提取Country
mydata$Country <- sub(".*/DONNEES/([^-]+)-.*", "\\1", mydata$path)

# 提取League
mydata$League <- sub(".*/DONNEES/[^-]+-(.*)-matches.*", "\\1", mydata$path)

# 提取Year
mydata$Year <- sub(".*-(\\d+)-to.*", "\\1", mydata$path)

# 查看最终结果
print(mydata)

正则说明:

  • Country:.*DONNEES/([^-]+)-.* 匹配到DONNEES/后,捕获第一个-前的所有非-字符(国家名)
  • League:.*DONNEES/[^-]+-(.*)-matches.* 跳过国家名部分,捕获到-matches前的内容(联赛名)
  • Year:.*-(\\d+)-to.* 捕获-to前的数字序列(年份)

方法2:使用stringr包(更直观)

stringr的str_extract函数支持正则预查,代码可读性更强:

library(stringr)

# 提取Country
mydata$Country <- str_extract(mydata$path, "(?<=DONNEES/)[^-]+")

# 提取League
mydata$League <- str_extract(mydata$path, "(?<=DONNEES/[^-]+-).*(?=-matches)")

# 提取Year
mydata$Year <- str_extract(mydata$path, "\\d+(?=-to)")

# 查看最终结果
print(mydata)

正则说明:

  • (?<=DONNEES/):正向预查,确保匹配内容前是DONNEES/
  • (?=-matches):正向预查,确保匹配内容后是-matches
  • \\d+(?=-to):匹配-to前的所有数字

内容的提问来源于stack exchange,提问作者Seydou GORO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 08:11:28