You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中从文本字符串提取体温相关句子及数值

R语言提取核心体温句子及数值的解决方案

原始数据集

df <- data.frame(
  Person.ID = c(123L, 234L),
  Date = c("10/10/09", "11/11/03"),
  Text = c(
    "Here are some random words that I do not want. The person was allowed to cool to a core body temperature of 16.5 degrees centigrade. Here are some other random words I do not want.",
    "Here are some random words that I do not want. A cooling mechanism was applied to cool the patient to a core body temperature of 19.1 degrees centigrade. Here are some other random words I do not want."
  )
)

需求说明

为每个个体提取提及核心体温的句子(去除无关内容),同时新增仅包含体温数值的列,预期输出如下:

df2 <- data.frame(
  Person.ID = c(123L, 234L),
  Date = c("10/10/09", "11/11/03"),
  Text = c(
    "The person was allowed to cool to a core body temperature of 16.5 degrees centigrade.",
    "A cooling mechanism was applied to cool the patient to a core body temperature of 19.1 degrees centigrade. "
  ),
  Value = c(16.5, 19.1)
)

实现代码

用tidyverse工具集就能搞定,代码如下:

# 加载所需包
library(tidyverse)

# 处理数据
df2 <- df %>%
  # 提取包含核心体温的完整句子
  mutate(Text = str_extract(Text, "[^.]*core body temperature[^.]*\\.")) %>%
  # 提取体温数值并转为数值型
  mutate(Value = as.numeric(str_extract(Text, "\\d+\\.\\d+")))

代码解释

  • str_extract(Text, "[^.]*core body temperature[^.]*\\."):匹配从任意非句号字符开始、包含"core body temperature"、到下一个句号结束的内容,精准提取目标句子。
  • str_extract(Text, "\\d+\\.\\d+"):匹配文本中的小数数值,再用as.numeric转换为数值类型,得到单独的体温数值列。

内容的提问来源于stack exchange,提问作者Jamie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 10:53:11