R data.table:如何基于列间字符串匹配正确筛选行?
data.table Filtering Questions Hey there! Let's work through your two data.table filtering issues—sounds like you're hitting a problem where only the first row is returned (with a warning) instead of all matching rows. Let's break it down step by step.
1. Using grepl in data.table's i parameter to match against another column
The key here is leveraging vectorization—grepl accepts vector inputs for both the pattern and x arguments, which aligns perfectly with data.table's column-wise operations.
First, let's set up a sample dataset to demonstrate:
library(data.table) dt <- data.table( fruit = c("apple", "banana", "cherry", "date"), description = c("Crunchy apple snack", "Banana bread recipe", "Berry smoothie", "Date cookie") )
To filter rows where description contains the string from the corresponding fruit column, use grepl directly in the i parameter:
# Case-sensitive match dt[grepl(fruit, description)] # Case-insensitive match (useful if capitalization varies) dt[grepl(fruit, description, ignore.case = TRUE)]
This works because grepl will pair each element in fruit with the corresponding element in description, returning a logical vector that data.table uses to select matching rows.
2. Filtering rows where column A's string exists in column B
This is exactly the same operation as the first question! The goal is to check, row-by-row, if the string in column A is present in column B. Using the same sample data, the code above already does this—it returns rows 1, 2, and 4 (since "cherry" isn't in "Berry smoothie").
Why were you only getting the first row + a warning?
Chances are you accidentally used a scalar value instead of the full column vector in grepl. For example, if you wrote:
# ❌ Wrong: Uses only the first element of `fruit` as the pattern dt[grepl(fruit[1], description)]
This would only find rows where description contains "apple" (the first value in fruit), hence only returning the first row. If you had a mismatch in vector lengths (e.g., passing a single value for x instead of the full column), data.table would warn about recycling values, leading to unexpected results.
Alternative: Using stringr::str_detect
If you prefer the stringr syntax, it works just as well in data.table:
library(stringr) dt[str_detect(description, fruit)]
内容的提问来源于stack exchange,提问作者moman822

