Data Frame列文本匹配求助:Q1关键词匹配无结果排查
Hey there! Let's break down why your current code isn't returning matches and fix it up.
The Core Issues
Your approach has two key hurdles right now:
- You're using the entire
keywordscolumn to match every row'sQ1instead of using each row's own keywords to check its correspondingQ1. That's not aligned with your goal of matching per-row keywords to per-row research questions. - Your
keywordscolumn is a single string of space-separated terms, not a split list of individual keywords. When you just concatenate them with|, you're creating a regex that looks for full phrases (like the entire first row's keywords as one match) instead of individual terms.
The Fix
We'll adjust the workflow to split keywords per row, then match each row's keywords against its own Q1 using row-wise processing. Here's the revised code:
library(tidyverse) # Load your sample data df <- structure( list( Q1 = c( "Assessing the effects of strategic deterrence messaging in the cognitive dimension", "How do you assess effects of strategic deterrence messaging?", "Determine Strategic Implications of Climate Change to USG/DoD" ), keywords = c( "Deterrence messaging effects perception assessment", "political philosophy sociology social sciences history marketing power structure government governing class bourgeoisie social class military class ruling class governing class", "Climate Change Strategic Global Warming Strategic Climate Change Policy Global Warming Policy" ) ), .Names = c("Q1", "keywords"), row.names = c(NA, -3L), class = c("tbl_df", "tbl", "data.frame") ) # Process the data to match keywords and count occurrences df_final <- df %>% # Split each row's keywords into a list of individual terms, removing duplicates mutate(keyword_list = str_split(keywords, "\\s+") %>% map(unique)) %>% # Switch to row-wise processing so we handle each row independently rowwise() %>% mutate( # Extract all matching keywords from Q1 (case-insensitive) matches = str_extract_all(Q1, regex(str_c(keyword_list, collapse = "|"), ignore_case = TRUE)), # Collapse matched terms into a single string for readability match = str_c(unlist(matches), collapse = ", "), # Count the total number of matches count = length(unlist(matches)) ) %>% # Exit row-wise mode to return to normal dataframe operations ungroup() # View the result df_final
What This Does
str_split(keywords, "\\s+") %>% map(unique): Splits each row's space-separated keywords into a vector, and removes duplicate terms (like the repeated "Strategic" in the third row).rowwise(): Ensures that every subsequent operation only uses data from the current row—so we're matching row 1's keywords to row 1's Q1, not the entire dataset's keywords.str_extract_all(...): Pulls all case-insensitive matches from Q1 using the row's specific keywords.- The final
matchandcountcolumns give you a human-readable list of matches and the total number of hits per row.
Sample Output
When you run this, you'll get results like:
- Row 1: Matches
effects, deterrence, messagingwith a count of 3 - Row 2: Matches
effects, deterrence, messagingwith a count of 3 - Row 3: Matches
Strategic, Climate Changewith a count of 2
If you actually need to match multi-word phrases (like treating "Climate Change" as a single term instead of two separate keywords), let me know—we'd need to adjust how we split the keywords column (e.g., if phrases are separated by a specific delimiter like commas).
内容的提问来源于stack exchange,提问作者Jason Heppler
相关产品推荐
相关产品推荐

