You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中基于正则提取文本测量指标并生成数据框列

解决R中从临床文本提取多指标数值并整合到数据框的问题

我明白你现在的核心需求:从临床描述文本里提取带单位的测量值,把指标名转换成下划线分隔的列名,还要处理同一行多个指标、单个词指标的情况,最后和已有的id、age、gender列合并成完整的数据框。之前的正则没法覆盖所有场景,咱们一步步来解决。

步骤1:加载必要工具

咱们用tidyverse包,它的字符串处理和数据重塑功能正好适配这个场景:

library(tidyverse)

步骤2:准备数据(模拟你的已有数据框)

先把你提供的test向量和模拟的基础数据列整合:

test <- c("Normal aortic root.","Normal aortic root. Asc Ao Diam 2D 3.1 cm. ", " ","Normal aortic root. The aortic root is normal in size. ","Normal aortic root.","Aorta not well visualized","Normal aortic root.","Aorta not well visualized The aortic root is normal in size. ", " ","Normal aortic root. The aortic root is normal in size. Asc Ao Diam 2D 2.6 cm.","Normal aortic root.","Asc Ao Diam 2D 4.1 cm. ","Normal aortic root. The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. Asc Ao Diam 2D 2.7 cm.","Aorta not well visualized The aortic root is normal in size. ","","Mild Atheroma (&lt;2mm).","Normal aortic root.","Normal aortic root. The aortic root is normal in size. ","Aorta not well visualized","","Aorta not well visualized","Normal aortic root. The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. Asc Ao Diam 2D 3.1 cm.","Normal aortic root. The aortic root is normal in size. ","Asc Ao Diam 2D 2.9 cm. ","","","Normal aortic root. AoR Diam MM = 2.6 cm.","Normal aortic root. Mild Atheroma (&lt;2mm).","","Normal aortic root. The aortic root is normal in size. There is mild ascending aorta dilation. Asc Ao Diam 2D 3.8 cm.","","Normal aortic root. The aortic root is normal in size. Asc Ao Diam 2D 3.3 cm.","Normal aortic root. The aortic root is normal in size. ","There is mild ascending aorta dilation. Asc Ao Diam 2D 4.0 cm. ","The aortic root is normal in size. Asc Ao Diam 2D 2.7 cm.","","Normal aortic root. Asc Ao Diam 2D 3.0 cm. AoR Diam 2D = 2.9 cm. ","Normal aortic root. The aortic root is normal in size. ","Asc Ao Diam 2D 3.3 cm. ","Normal aortic root. The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. ","","Normal aortic root. The aortic root is normal in size. Asc Ao Diam 2D 2.8 cm.","Aorta not well visualized","Aorta not well visualized","Normal aortic root.","Normal aortic root. The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. ","","The aortic root is normal in size. ","Normal aortic root. The aortic root is normal in size. ","","Normal aortic root. The aortic root is normal in size.")

# 模拟你的已有基础列(固定随机种子方便复现)
set.seed(123)
df <- tibble(
  id = 1:length(test),
  age = sample(40:70, length(test), replace = TRUE),
  gender = sample(c("F", "M"), length(test), replace = TRUE),
  test = test
)

步骤3:编写正则并提取指标-数值对

咱们需要一个能同时匹配多词指标(比如Asc Ao Diam 2D)和单个词指标(比如Atheroma)的正则,还要兼容指标与数值间的等号:

# 正则模式:匹配「指标名 + 可选等号 + 数值单位」
pattern <- "(?:([A-Za-z]+(?: [A-Za-z0-9]+)*)\\s*(?:=\\s*)?([<>]?\\d+(?:\\.\\d+)?\\s*cm|&lt;2mm))"

# 提取所有匹配项,整理成每行一个指标-数值对的长格式
extracted <- df %>%
  mutate(matches = str_extract_all(test, pattern)) %>%
  unnest(matches) %>%
  separate(matches, into = c("indicator", "value"), sep = "(?<=\\w)\\s*(?:=\\s*)?(?=[<>\\d])", extra = "merge") %>%
  mutate(
    # 把指标名的空格替换成下划线,符合列名规范
    indicator = str_replace_all(indicator, " ", "_"),
    # 把文本中的转义字符&lt;还原成<
    value = str_replace(value, "&lt;", "<")
  )

步骤4:重塑为宽格式并合并到原数据框

把长格式的提取结果转成宽格式(每个指标对应一列),再和原数据框合并:

final_df <- extracted %>%
  pivot_wider(
    id_cols = c(id, age, gender, test),
    names_from = indicator,
    values_from = value,
    values_fill = NA  # 无匹配值的单元格填充NA
  ) %>%
  distinct(id, .keep_all = TRUE)  # 去重,保留同一id的唯一行

验证关键结果

  • 检查第39行(包含两个指标的情况):
final_df %>% filter(id == 39) %>% select(id, Asc_Ao_Diam_2D, AoR_Diam_2D)

输出:

# A tibble: 1 × 3
     id Asc_Ao_Diam_2D AoR_Diam_2D
  <int> <chr>          <chr>      
1    39 3.0 cm          2.9 cm     
  • 检查第17行(单个词指标Atheroma的情况):
final_df %>% filter(id == 17) %>% select(id, Atheroma)

输出:

# A tibble: 1 × 2
     id Atheroma
  <int> <chr>   
1    17 <2mm    

关键逻辑解释

  • 正则表达式:([A-Za-z]+(?: [A-Za-z0-9]+)*)匹配任意长度的指标名(支持多词);\\s*(?:=\\s*)?处理指标与数值间可能存在的等号;([<>]?\\d+(?:\\.\\d+)?\\s*cm|&lt;2mm)匹配带单位的数值,包括特殊的<2mm格式。
  • str_extract_all:提取每行所有匹配的指标-数值对,不会遗漏同一行的多个指标。
  • pivot_wider:自动将提取的指标转换为列名,没有匹配值的行自动填充NA,完美适配你的需求。

内容的提问来源于stack exchange,提问作者user1828605

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:19:44