You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas DataFrame生成含非匹配短语样本的新列技术需求

Solution for Adding Random Valid Phrases Column to DataFrame

Got it, let's work through this problem step by step. I'll show you how to create that new column in R, using your provided DataFrame structure and a sample phrase list.

Step 1: Define Your Input Data

First, let's recreate your DataFrame and set up a sample phrase list (you can replace this with your actual list):

# Load required library (dplyr for easy row-wise operations)
library(dplyr)

# Your original DataFrame
df <- structure(
  list(report = c(
    "Biopsies of small bowel mucosa including Brunner's glands",
    "These are fragments of small bowel mucosa which include Brunner's glands ",
    "These are fragments of small bowel mucosa which include..."
  )),
  class = "data.frame",
  row.names = c(NA, -3L)
)

# Sample phrase list (replace with your actual phrases)
phrase_list <- c("Brunner's glands", "colon mucosa", "gastric biopsy", "liver parenchyma", "pancreatic tissue")

Step 2: Create a Custom Function to Get Valid Random Phrases

Next, we'll write a function that takes a single report text and the phrase list, then returns a random number of phrases that don't appear in the report. We'll let the random count range from 1 to 3 (you can adjust this by changing the min_n and max_n values):

get_random_valid_phrases <- function(report_text, phrase_list, min_n = 1, max_n = 3) {
  # Filter phrases that are NOT present in the report text (uses word boundaries to avoid partial matches)
  valid_phrases <- phrase_list[!grepl(paste0("\\b", phrase_list, "\\b"), report_text, ignore.case = TRUE)]
  
  # Handle edge case where no valid phrases are left
  if (length(valid_phrases) == 0) {
    return(NA_character_)
  }
  
  # Randomly select between min_n and max_n phrases (capped at available valid phrases)
  sample_size <- sample(min(min_n, length(valid_phrases)):min(max_n, length(valid_phrases)), 1)
  selected_phrases <- sample(valid_phrases, sample_size)
  
  # Join selected phrases into a single string (adjust separator as needed)
  paste(selected_phrases, collapse = ", ")
}

Step 3: Apply the Function to Your DataFrame

Now we'll use dplyr::rowwise() to apply the function to each row, creating our new column (let's call it random_valid_phrases):

df_with_new_col <- df %>%
  rowwise() %>%
  mutate(random_valid_phrases = get_random_valid_phrases(report, phrase_list)) %>%
  ungroup()

# View the result
print(df_with_new_col)

Key Details to Note

  • The grepl uses word boundaries (\\b) to avoid partial matches (e.g., it won't mistake "Brunner's" for "Brunner's glands").
  • We added a check for when all phrases are present in a report—this returns NA (you can change this to an empty string or custom message if needed).
  • The sample size is capped at the number of valid phrases, so we never try to select more phrases than are available.

Example output might look like this (since it's random, your results will vary):

A tibble: 3 × 2

report random_valid_phrases

1 Biopsies of small bowel mucosa including Brunner's glands gastric biopsy, liver parenchyma
2 These are fragments of small bowel mucosa which include Brunner's glands pancreatic tissue, colon mucosa
3 These are fragments of small bowel mucosa which include... Brunner's glands, gastric biopsy

内容的提问来源于stack exchange,提问作者Sebastian Zeki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:37:20