You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用stringr处理非结构化文本并计算药物周期总剂量?

Got it, let’s walk through how to tackle both your dose calculation and unstructured text processing needs using R—specifically combining dplyr for data manipulation and stringr for wrangling that messy EXPANDED_INSTRUCTIONS field.

1. Calculating Total Daily/Weekly/Biweekly/Monthly Drug Doses

The core formula for total dose is straightforward:
Total Dose = Single Dose Strength (DRUG_STRENGTH_NO) × Doses Per Day × Days in Period

First, we need to extract the number of doses per day from EXPANDED_INSTRUCTIONS, since that’s the unstructured piece. The exact regex will depend on whether your instructions are in English or Chinese, but here are examples for both scenarios:

Example for English Instructions

library(dplyr)
library(stringr)

# Start with your data frame
df <- your_data_frame_name

# Extract daily dose count and calculate totals
df <- df %>%
  mutate(
    # Pull number of doses per day from text
    daily_dose_count = case_when(
      # Match common abbreviations or phrases
      str_detect(EXPANDED_INSTRUCTIONS, "\\bqd\\b|once daily") ~ 1,
      str_detect(EXPANDED_INSTRUCTIONS, "\\bbid\\b|twice daily") ~ 2,
      str_detect(EXPANDED_INSTRUCTIONS, "\\btid\\b|three times daily") ~ 3,
      str_detect(EXPANDED_INSTRUCTIONS, "\\bqid\\b|four times daily") ~ 4,
      # Extract numeric values from phrases like "3 times daily"
      str_detect(EXPANDED_INSTRUCTIONS, "\\d+ times daily") ~ as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "\\d+")),
      # Convert hourly instructions to daily count (e.g., every 8 hours = 3 doses/day)
      str_detect(EXPANDED_INSTRUCTIONS, "every \\d+ hours") ~ 24 / as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "(?<=every )\\d+")),
      # Fallback if no match is found
      TRUE ~ NA_real_
    ),
    # Calculate daily total dose
    daily_total = DRUG_STRENGTH_NO * daily_dose_count,
    # Extend to other periods
    weekly_total = daily_total * 7,
    biweekly_total = daily_total * 14,
    monthly_total = daily_total * 30.44  # Use average monthly days for consistency
  )

Example for Chinese Instructions

If your EXPANDED_INSTRUCTIONS are in Chinese (e.g., "每日两次,每次50mg"), adjust the regex to match Chinese keywords:

df <- df %>%
  mutate(
    daily_dose_count = case_when(
      str_detect(EXPANDED_INSTRUCTIONS, "每日一次|qd") ~ 1,
      str_detect(EXPANDED_INSTRUCTIONS, "每日两次|bid") ~ 2,
      str_detect(EXPANDED_INSTRUCTIONS, "每日\\d+次") ~ as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "\\d+")),
      str_detect(EXPANDED_INSTRUCTIONS, "每\\d+小时一次") ~ 24 / as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "(?<=每)\\d+")),
      TRUE ~ NA_real_
    ),
    # Same calculations as above
    daily_total = DRUG_STRENGTH_NO * daily_dose_count,
    weekly_total = daily_total * 7,
    biweekly_total = daily_total * 14,
    monthly_total = daily_total * 30.44
  )
2. Using stringr to Process Unstructured Text

stringr is perfect for cleaning and extracting info from messy text. Here are the most useful functions for your use case, with examples:

  • str_detect(): Check if text contains a specific pattern (used earlier to identify dose frequency)

    # Flag rows with daily dosing instructions
    df <- df %>%
      mutate(is_daily = str_detect(EXPANDED_INSTRUCTIONS, "daily|每日"))
    
  • str_extract()/str_extract_all(): Pull specific values from text (e.g., extract numeric doses or time intervals)

    # Extract any numeric value that might represent a dose adjustment
    df <- df %>%
      mutate(adjustment_dose = as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "\\d+\\s?(mg|g)")))
    
  • str_replace_all(): Clean up messy text (remove punctuation, standardize abbreviations)

    # Remove extra punctuation and convert to lowercase (for English text)
    df <- df %>%
      mutate(cleaned_instructions = str_to_lower(str_replace_all(EXPANDED_INSTRUCTIONS, "[^a-zA-Z0-9\\s]", "")))
    
    # For Chinese text: remove special characters
    df <- df %>%
      mutate(cleaned_instructions = str_replace_all(EXPANDED_INSTRUCTIONS, "[^\\u4e00-\\u9fa50-9\\s]", ""))
    
  • str_split(): Split long instructions into smaller chunks (useful for complex orders)

    # Split instructions by commas into a list column
    df <- df %>%
      mutate(instruction_chunks = str_split(EXPANDED_INSTRUCTIONS, ","))
    

Pro Tip

Non-structured text is messy—you’ll likely need to tweak your regex patterns to match the specific language and formatting in your data. Test small subsets first to make sure your patterns are capturing all cases, and add more case_when() conditions as you find edge cases.

内容的提问来源于stack exchange,提问作者HNSKD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:34:50