You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Data Frame列文本匹配求助:Q1关键词匹配无结果排查

Hey there! Let's break down why your current code isn't returning matches and fix it up.

The Core Issues

Your approach has two key hurdles right now:

  1. You're using the entire keywords column to match every row's Q1 instead of using each row's own keywords to check its corresponding Q1. That's not aligned with your goal of matching per-row keywords to per-row research questions.
  2. Your keywords column is a single string of space-separated terms, not a split list of individual keywords. When you just concatenate them with |, you're creating a regex that looks for full phrases (like the entire first row's keywords as one match) instead of individual terms.

The Fix

We'll adjust the workflow to split keywords per row, then match each row's keywords against its own Q1 using row-wise processing. Here's the revised code:

library(tidyverse)

# Load your sample data
df <- structure(
  list(
    Q1 = c(
      "Assessing the effects of strategic deterrence messaging in the cognitive dimension",
      "How do you assess effects of strategic deterrence messaging?",
      "Determine Strategic Implications of Climate Change to USG/DoD"
    ),
    keywords = c(
      "Deterrence messaging effects perception assessment",
      "political philosophy sociology social sciences history marketing power structure government governing class bourgeoisie social class military class ruling class governing class",
      "Climate Change Strategic Global Warming Strategic Climate Change Policy Global Warming Policy"
    )
  ),
  .Names = c("Q1", "keywords"),
  row.names = c(NA, -3L),
  class = c("tbl_df", "tbl", "data.frame")
)

# Process the data to match keywords and count occurrences
df_final <- df %>%
  # Split each row's keywords into a list of individual terms, removing duplicates
  mutate(keyword_list = str_split(keywords, "\\s+") %>% map(unique)) %>%
  # Switch to row-wise processing so we handle each row independently
  rowwise() %>%
  mutate(
    # Extract all matching keywords from Q1 (case-insensitive)
    matches = str_extract_all(Q1, regex(str_c(keyword_list, collapse = "|"), ignore_case = TRUE)),
    # Collapse matched terms into a single string for readability
    match = str_c(unlist(matches), collapse = ", "),
    # Count the total number of matches
    count = length(unlist(matches))
  ) %>%
  # Exit row-wise mode to return to normal dataframe operations
  ungroup()

# View the result
df_final

What This Does

  • str_split(keywords, "\\s+") %>% map(unique): Splits each row's space-separated keywords into a vector, and removes duplicate terms (like the repeated "Strategic" in the third row).
  • rowwise(): Ensures that every subsequent operation only uses data from the current row—so we're matching row 1's keywords to row 1's Q1, not the entire dataset's keywords.
  • str_extract_all(...): Pulls all case-insensitive matches from Q1 using the row's specific keywords.
  • The final match and count columns give you a human-readable list of matches and the total number of hits per row.

Sample Output

When you run this, you'll get results like:

  • Row 1: Matches effects, deterrence, messaging with a count of 3
  • Row 2: Matches effects, deterrence, messaging with a count of 3
  • Row 3: Matches Strategic, Climate Change with a count of 2

If you actually need to match multi-word phrases (like treating "Climate Change" as a single term instead of two separate keywords), let me know—we'd need to adjust how we split the keywords column (e.g., if phrases are separated by a specific delimiter like commas).

内容的提问来源于stack exchange,提问作者Jason Heppler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:23:16