基于Pandas DataFrame生成含非匹配短语样本的新列技术需求
Got it, let's work through this problem step by step. I'll show you how to create that new column in R, using your provided DataFrame structure and a sample phrase list.
Step 1: Define Your Input Data
First, let's recreate your DataFrame and set up a sample phrase list (you can replace this with your actual list):
# Load required library (dplyr for easy row-wise operations) library(dplyr) # Your original DataFrame df <- structure( list(report = c( "Biopsies of small bowel mucosa including Brunner's glands", "These are fragments of small bowel mucosa which include Brunner's glands ", "These are fragments of small bowel mucosa which include..." )), class = "data.frame", row.names = c(NA, -3L) ) # Sample phrase list (replace with your actual phrases) phrase_list <- c("Brunner's glands", "colon mucosa", "gastric biopsy", "liver parenchyma", "pancreatic tissue")
Step 2: Create a Custom Function to Get Valid Random Phrases
Next, we'll write a function that takes a single report text and the phrase list, then returns a random number of phrases that don't appear in the report. We'll let the random count range from 1 to 3 (you can adjust this by changing the min_n and max_n values):
get_random_valid_phrases <- function(report_text, phrase_list, min_n = 1, max_n = 3) { # Filter phrases that are NOT present in the report text (uses word boundaries to avoid partial matches) valid_phrases <- phrase_list[!grepl(paste0("\\b", phrase_list, "\\b"), report_text, ignore.case = TRUE)] # Handle edge case where no valid phrases are left if (length(valid_phrases) == 0) { return(NA_character_) } # Randomly select between min_n and max_n phrases (capped at available valid phrases) sample_size <- sample(min(min_n, length(valid_phrases)):min(max_n, length(valid_phrases)), 1) selected_phrases <- sample(valid_phrases, sample_size) # Join selected phrases into a single string (adjust separator as needed) paste(selected_phrases, collapse = ", ") }
Step 3: Apply the Function to Your DataFrame
Now we'll use dplyr::rowwise() to apply the function to each row, creating our new column (let's call it random_valid_phrases):
df_with_new_col <- df %>% rowwise() %>% mutate(random_valid_phrases = get_random_valid_phrases(report, phrase_list)) %>% ungroup() # View the result print(df_with_new_col)
Key Details to Note
- The
grepluses word boundaries (\\b) to avoid partial matches (e.g., it won't mistake "Brunner's" for "Brunner's glands"). - We added a check for when all phrases are present in a report—this returns
NA(you can change this to an empty string or custom message if needed). - The sample size is capped at the number of valid phrases, so we never try to select more phrases than are available.
Example output might look like this (since it's random, your results will vary):
A tibble: 3 × 2report random_valid_phrases
1 Biopsies of small bowel mucosa including Brunner's glands gastric biopsy, liver parenchyma
2 These are fragments of small bowel mucosa which include Brunner's glands pancreatic tissue, colon mucosa
3 These are fragments of small bowel mucosa which include... Brunner's glands, gastric biopsy
内容的提问来源于stack exchange,提问作者Sebastian Zeki

