如何用R语言gsub提取成分列表?移除句首/标点至冒号间内容
gsub() Hey there! Let's solve this problem using R's gsub() function to clean up your ingredient string. Your goal is to strip out all the label prefixes and combine the ingredients into a single list—here are two straightforward approaches to get you there:
Approach 1: One-Step Regex Replacement
This method uses a single gsub() call to handle both the initial label and the middle separator in one go:
# Your input string input_str <- "active ingredients: avobenzone, octocrylene, octyl salicylate. other stuff inactive ingredients: water, glycerin, edta." # Apply the regex replacement cleaned_ingredients <- gsub("^.*?:|\\s*\\.\\s*[^:]+?:", ", ", input_str) # Check the result cat(cleaned_ingredients) # Output: "avobenzone, octocrylene, octyl salicylate, water, glycerin, edta."
Regex Breakdown:
^.*?:: Matches everything from the start of the string up to the first colon (:), and replaces it with nothing. This removes the initial "active ingredients:" label.\\s*\\.\\s*[^:]+?:: Matches the period (.) plus any surrounding spaces, followed by all characters up to the next colon (this targets the ". other stuff inactive ingredients:" part). We replace this with", "to smoothly connect the two ingredient lists.
Approach 2: Step-by-Step Cleanup (More Readable for Beginners)
If you prefer breaking the process into simpler steps, this method is easier to follow and debug:
input_str <- "active ingredients: avobenzone, octocrylene, octyl salicylate. other stuff inactive ingredients: water, glycerin, edta." # Step 1: Remove all label prefixes (everything up to each colon) step1 <- gsub(".*?:", "", input_str) # Step 1 result: " avobenzone, octocrylene, octyl salicylate. water, glycerin, edta." # Step 2: Replace the period and surrounding spaces with a comma and space step2 <- gsub("\\.\\s*", ", ", step1) # Step 2 result: " avobenzone, octocrylene, octyl salicylate, water, glycerin, edta." # Step 3: Trim leading/trailing whitespace cleaned_ingredients <- trimws(step2) # Check the result cat(cleaned_ingredients) # Output: "avobenzone, octocrylene, octyl salicylate, water, glycerin, edta."
Both methods will give you exactly the cleaned ingredient list you're looking for. The one-step approach is concise, while the step-by-step version is great if you want to understand each part of the cleaning process.
内容的提问来源于stack exchange,提问作者sir_chocolate_soup

