使用Quanteda匹配医学多词短语并生成指定DataFrame求助
Fixing Multi-Word Phrase Matching with Quanteda for Medical Terms
Hey there! I see you're having trouble getting quanteda to match multi-word medical phrases (1-5 words) instead of just single terms. Let's fix that up—here's the adjusted code and a breakdown of what was missing:
Full Working Code
library(quanteda) library(dplyr) library(tidyr) # Your original raw text data raw <- data.frame( "doc_id" = c("1", "2", "3"), "text" = c( "diffuse intrinsic pontine glioma are highly aggressive and difficult to treat brain tumors found at the base of the brain.", "magnetic resonance imaging (mri) is a medical imaging technique used in radiology to form pictures of the anatomy and the physiological processes of the body.", "radiation therapy or radiotherapy, often abbreviated rt, rtx, or xrt, is a therapy using ionizing radiation, generally as part of cancer treatment to control or kill malignant cells and normally delivered by a linear accelerator." ) ) # Your medical term list converted to a quanteda dictionary term_list <- c( "diffuse intrinsic pontine glioma", "brain tumors", "brain", "pontine glioma", "mri", "medical imaging", "radiology", "anatomy", "physiological processes", "radiation therapy", "radiotherapy", "cancer treatment", "malignant cells" ) med_dict <- dictionary(list(medical_terms = term_list)) # Create the corpus directly from your raw data frame corp <- corpus(raw, text_field = "text") # Generate a DFM with 1-5 word n-grams (this is the key fix!) multi_gram_dfm <- dfm( corp, tolower = TRUE, stem = FALSE, remove_punct = TRUE, remove = stopwords("english"), ngrams = 1:5, # Generate all 1 to 5-word combinations concatenator = " " # Join n-gram words with spaces (matches your dict format) ) # Match your medical terms against the multi-gram DFM matched_dfm <- dfm_select(multi_gram_dfm, pattern = phrase(med_dict)) # Convert to your desired DataFrame format result_df <- convert(matched_dfm, to = "data.frame") %>% pivot_longer(-doc_id, names_to = "phrase", values_to = "count") %>% filter(count > 0) %>% # Keep only phrases that actually appear in the text select(doc_id, phrase) %>% arrange(doc_id, phrase) # View the final result print(result_df)
Key Fixes Explained
ngrams = 1:5: This is the critical missing piece! Quanteda defaults to only single-word tokens. By specifying 1 to 5-word n-grams, we generate all possible word combinations in that length range, which allows us to match your multi-word medical phrases.concatenator = " ": This ensures that n-grams are joined with spaces (e.g., "diffuse intrinsic pontine glioma" instead of "diffuse_intrinsic_pontine_glioma"), which matches the format of the phrases in your dictionary. Without this, the matching would fail for multi-word terms.- DataFrame Conversion: We use
convert()to turn the DFM into a data frame, thenpivot_longer()to reshape it into the two-column format you want. Filtering out counts of 0 removes any terms that didn't appear in the document.
Expected Output
Running this code will give you exactly the result you were hoping for:
doc_id phrase 1 1 brain 2 1 brain tumors 3 1 diffuse intrinsic pontine glioma 4 1 pontine glioma 5 2 anatomy 6 2 medical imaging 7 2 mri 8 2 radiology 9 2 physiological processes 10 3 cancer treatment 11 3 malignant cells 12 3 radiation therapy 13 3 radiotherapy
内容的提问来源于stack exchange,提问作者Obed
相关产品推荐
相关产品推荐

