基于dplyr结构的文献数据库:从tag列提取语言至language列
tag Column to Populate language in dplyr Hey there! Let's work through how to extract language info from that deprecated tag column and fill your language column using dplyr (plus stringr for text handling—perfect combo for this task). Here's what you can do:
First, make sure you have the necessary packages loaded:
library(dplyr) library(stringr)
Step 1: Define your target languages
First, list out all the language terms you expect to find in the tag column (adjust this to match your actual data):
target_languages <- c("English", "Spanish", "French", "German", "Chinese", "Japanese")
Method 1: Split tags to validate matches (more thorough)
This approach breaks down the comma-separated tags into individual rows, filters only the language entries, then joins back to your original data. Great if you want to double-check which tags are being pulled as languages:
# Extract valid language tags from the tag column language_matches <- your_lit_db %>% select(book_id, tag) %>% # Keep only unique identifier and tag column separate_rows(tag, sep = ",\\s*") %>% # Split tags into rows (handles commas with/without spaces) filter(str_detect(tag, paste(target_languages, collapse = "|"))) %>% # Keep only rows with language matches rename(language_extracted = tag) # Join back to original data and populate the language column updated_lit_db <- your_lit_db %>% left_join(language_matches, by = "book_id") %>% # Keep existing language values if they're not NA; use extracted values otherwise mutate(language = coalesce(language, language_extracted)) %>% select(-language_extracted) # Clean up temporary column
Method 2: Direct extraction (faster, for clean data)
If you're confident your language terms don't overlap with other tags, you can extract directly in the original data frame without splitting rows:
updated_lit_db <- your_lit_db %>% mutate(language = coalesce( language, # Keep existing non-NA language values str_extract(tag, paste(target_languages, collapse = "|")) # Extract first matching language from tag ))
Handling edge cases
- Case inconsistency: If tags have mixed case (e.g., "english" vs "English"), adjust the regex to ignore case:
# For Method 1 filter filter(str_detect(tag, regex(paste(target_languages, collapse = "|"), ignore_case = TRUE))) # For Method 2 extraction str_extract(tag, regex(paste(target_languages, collapse = "|"), ignore_case = TRUE)) - Multiple languages per book: If a book has multiple language tags, use
str_extract_allto pull all matches and combine them:updated_lit_db <- your_lit_db %>% mutate(language = coalesce( language, str_extract_all(tag, paste(target_languages, collapse = "|")) %>% map_chr(~str_c(., collapse = ", ")) # Combine multiple languages into one string ))
Just replace your_lit_db with the name of your actual data frame, and tweak target_languages to match the language terms in your tag column!
内容的提问来源于stack exchange,提问作者miri sueß

