R语言中提取自由文本数字及对应年龄单位的多组信息需求
R语言提取自由文本中的多组年龄(数字+单位)解决方案
核心思路
用stringr包的str_match_all提取所有匹配的年龄组合(数字+单位),通过正则捕获组拆分数字和单位,再整理成结构化数据;用tidyr的unnest可将多组年龄展开为多行,方便后续处理。
完整代码示例
# 加载依赖包 library(dplyr) library(stringr) library(tidyr) # 示例数据集 dataset <- data.frame( Freetext = c( "A 3-year-old body was at the supermarket", "A 45 year old mother was at the supermarket with her little 5 months old boy", "The baby is 10 days-old and his grandma is 70 Jahre alt" ), Furtherinfo = c("ID1", "ID2", "ID3") ) # 处理年龄提取与结构化 dataset_processed <- dataset %>% mutate( # 提取所有匹配的年龄组,捕获数字和单位 age_matches = str_match_all(Freetext, "(\\d+)\\s*(?:(year|month|day)s?\\s*(?:old)?|-(year|month|day)-old|(year|month|day)s?-old|Jahre alt)") %>% lapply(function(x) { # 合并不同格式下的单位捕获组 units <- coalesce(x[,2], x[,3], x[,4]) # 统一德语单位为英文 units <- ifelse(units == "Jahre", "year", units) # 返回结构化的年龄数据框 data.frame( age_number = as.numeric(x[,1]), age_unit = units ) }) ) # 将多组年龄展开为多行(可选) dataset_unnested <- dataset_processed %>% unnest(age_matches, keep_empty = TRUE) # 查看结果 print(dataset_unnested)
代码说明
- 正则表达式:覆盖了常见的年龄格式(如
3-year-old、45 year old、5 months old、10 days-old、德语70 Jahre alt),通过捕获组分别提取数字和单位。 str_match_all:替代原代码的str_extract,支持提取文本中的多组年龄信息。- 单位统一:将德语的
Jahre转换为year,方便后续统一分析。 unnest:把每个文本对应的多组年龄从列表列展开为多行,适合需要逐行分析单年龄的场景。
输出结果示例
Freetext Furtherinfo age_number age_unit 1 A 3-year-old body was at the supermarket ID1 3 year 2 A 45 year old mother was at the supermarket with her little 5 months old boy ID2 45 year 3 A 45 year old mother was at the supermarket with her little 5 months old boy ID2 5 month 4 The baby is 10 days-old and his grandma is 70 Jahre alt ID3 10 day 5 The baby is 10 days-old and his grandma is 70 Jahre alt ID3 70 year
内容的提问来源于stack exchange,提问作者USER12345
相关产品推荐
相关产品推荐

