如何使用stringr处理非结构化文本并计算药物周期总剂量?
Got it, let’s walk through how to tackle both your dose calculation and unstructured text processing needs using R—specifically combining dplyr for data manipulation and stringr for wrangling that messy EXPANDED_INSTRUCTIONS field.
The core formula for total dose is straightforward:Total Dose = Single Dose Strength (DRUG_STRENGTH_NO) × Doses Per Day × Days in Period
First, we need to extract the number of doses per day from EXPANDED_INSTRUCTIONS, since that’s the unstructured piece. The exact regex will depend on whether your instructions are in English or Chinese, but here are examples for both scenarios:
Example for English Instructions
library(dplyr) library(stringr) # Start with your data frame df <- your_data_frame_name # Extract daily dose count and calculate totals df <- df %>% mutate( # Pull number of doses per day from text daily_dose_count = case_when( # Match common abbreviations or phrases str_detect(EXPANDED_INSTRUCTIONS, "\\bqd\\b|once daily") ~ 1, str_detect(EXPANDED_INSTRUCTIONS, "\\bbid\\b|twice daily") ~ 2, str_detect(EXPANDED_INSTRUCTIONS, "\\btid\\b|three times daily") ~ 3, str_detect(EXPANDED_INSTRUCTIONS, "\\bqid\\b|four times daily") ~ 4, # Extract numeric values from phrases like "3 times daily" str_detect(EXPANDED_INSTRUCTIONS, "\\d+ times daily") ~ as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "\\d+")), # Convert hourly instructions to daily count (e.g., every 8 hours = 3 doses/day) str_detect(EXPANDED_INSTRUCTIONS, "every \\d+ hours") ~ 24 / as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "(?<=every )\\d+")), # Fallback if no match is found TRUE ~ NA_real_ ), # Calculate daily total dose daily_total = DRUG_STRENGTH_NO * daily_dose_count, # Extend to other periods weekly_total = daily_total * 7, biweekly_total = daily_total * 14, monthly_total = daily_total * 30.44 # Use average monthly days for consistency )
Example for Chinese Instructions
If your EXPANDED_INSTRUCTIONS are in Chinese (e.g., "每日两次,每次50mg"), adjust the regex to match Chinese keywords:
df <- df %>% mutate( daily_dose_count = case_when( str_detect(EXPANDED_INSTRUCTIONS, "每日一次|qd") ~ 1, str_detect(EXPANDED_INSTRUCTIONS, "每日两次|bid") ~ 2, str_detect(EXPANDED_INSTRUCTIONS, "每日\\d+次") ~ as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "\\d+")), str_detect(EXPANDED_INSTRUCTIONS, "每\\d+小时一次") ~ 24 / as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "(?<=每)\\d+")), TRUE ~ NA_real_ ), # Same calculations as above daily_total = DRUG_STRENGTH_NO * daily_dose_count, weekly_total = daily_total * 7, biweekly_total = daily_total * 14, monthly_total = daily_total * 30.44 )
stringr to Process Unstructured Text stringr is perfect for cleaning and extracting info from messy text. Here are the most useful functions for your use case, with examples:
str_detect(): Check if text contains a specific pattern (used earlier to identify dose frequency)# Flag rows with daily dosing instructions df <- df %>% mutate(is_daily = str_detect(EXPANDED_INSTRUCTIONS, "daily|每日"))str_extract()/str_extract_all(): Pull specific values from text (e.g., extract numeric doses or time intervals)# Extract any numeric value that might represent a dose adjustment df <- df %>% mutate(adjustment_dose = as.numeric(str_extract(EXPANDED_INSTRUCTIONS, "\\d+\\s?(mg|g)")))str_replace_all(): Clean up messy text (remove punctuation, standardize abbreviations)# Remove extra punctuation and convert to lowercase (for English text) df <- df %>% mutate(cleaned_instructions = str_to_lower(str_replace_all(EXPANDED_INSTRUCTIONS, "[^a-zA-Z0-9\\s]", ""))) # For Chinese text: remove special characters df <- df %>% mutate(cleaned_instructions = str_replace_all(EXPANDED_INSTRUCTIONS, "[^\\u4e00-\\u9fa50-9\\s]", ""))str_split(): Split long instructions into smaller chunks (useful for complex orders)# Split instructions by commas into a list column df <- df %>% mutate(instruction_chunks = str_split(EXPANDED_INSTRUCTIONS, ","))
Pro Tip
Non-structured text is messy—you’ll likely need to tweak your regex patterns to match the specific language and formatting in your data. Test small subsets first to make sure your patterns are capturing all cases, and add more case_when() conditions as you find edge cases.
内容的提问来源于stack exchange,提问作者HNSKD

