You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Quanteda匹配医学多词短语并生成指定DataFrame求助

Fixing Multi-Word Phrase Matching with Quanteda for Medical Terms

Hey there! I see you're having trouble getting quanteda to match multi-word medical phrases (1-5 words) instead of just single terms. Let's fix that up—here's the adjusted code and a breakdown of what was missing:

Full Working Code

library(quanteda)
library(dplyr)
library(tidyr)

# Your original raw text data
raw <- data.frame(
  "doc_id" = c("1", "2", "3"), 
  "text" = c(
    "diffuse intrinsic pontine glioma are highly aggressive and difficult to treat brain tumors found at the base of the brain.", 
    "magnetic resonance imaging (mri) is a medical imaging technique used in radiology to form pictures of the anatomy and the physiological processes of the body.", 
    "radiation therapy or radiotherapy, often abbreviated rt, rtx, or xrt, is a therapy using ionizing radiation, generally as part of cancer treatment to control or kill malignant cells and normally delivered by a linear accelerator."
  )
)

# Your medical term list converted to a quanteda dictionary
term_list <- c(
  "diffuse intrinsic pontine glioma", "brain tumors", "brain", 
  "pontine glioma", "mri", "medical imaging", "radiology", 
  "anatomy", "physiological processes", "radiation therapy", 
  "radiotherapy", "cancer treatment", "malignant cells"
)
med_dict <- dictionary(list(medical_terms = term_list))

# Create the corpus directly from your raw data frame
corp <- corpus(raw, text_field = "text")

# Generate a DFM with 1-5 word n-grams (this is the key fix!)
multi_gram_dfm <- dfm(
  corp,
  tolower = TRUE,
  stem = FALSE,
  remove_punct = TRUE,
  remove = stopwords("english"),
  ngrams = 1:5,  # Generate all 1 to 5-word combinations
  concatenator = " "  # Join n-gram words with spaces (matches your dict format)
)

# Match your medical terms against the multi-gram DFM
matched_dfm <- dfm_select(multi_gram_dfm, pattern = phrase(med_dict))

# Convert to your desired DataFrame format
result_df <- convert(matched_dfm, to = "data.frame") %>%
  pivot_longer(-doc_id, names_to = "phrase", values_to = "count") %>%
  filter(count > 0) %>%  # Keep only phrases that actually appear in the text
  select(doc_id, phrase) %>%
  arrange(doc_id, phrase)

# View the final result
print(result_df)

Key Fixes Explained

  • ngrams = 1:5: This is the critical missing piece! Quanteda defaults to only single-word tokens. By specifying 1 to 5-word n-grams, we generate all possible word combinations in that length range, which allows us to match your multi-word medical phrases.
  • concatenator = " ": This ensures that n-grams are joined with spaces (e.g., "diffuse intrinsic pontine glioma" instead of "diffuse_intrinsic_pontine_glioma"), which matches the format of the phrases in your dictionary. Without this, the matching would fail for multi-word terms.
  • DataFrame Conversion: We use convert() to turn the DFM into a data frame, then pivot_longer() to reshape it into the two-column format you want. Filtering out counts of 0 removes any terms that didn't appear in the document.

Expected Output

Running this code will give you exactly the result you were hoping for:

doc_id                          phrase
1       1                           brain
2       1                   brain tumors
3       1 diffuse intrinsic pontine glioma
4       1                  pontine glioma
5       2                         anatomy
6       2                 medical imaging
7       2                             mri
8       2                       radiology
9       2            physiological processes
10      3                 cancer treatment
11      3                  malignant cells
12      3                radiation therapy
13      3                    radiotherapy

内容的提问来源于stack exchange,提问作者Obed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:42:48