为何str_detect无法直接接收列名参数?tidyverse使用疑问解答
关于tidyverse中str_detect使用的问题解答
示例数据
structure(list(Product_Name = c("Delicious Chips", "Creamy Tomato Soup", "Cheesy Macaroni", "Savory Meatballs", "Crispy Chicken Tenders" ), Ingredients = c("Potato Slices | Vegetable Oil | Salt | Seasoning Blend", "Tomatoes | Water | Cream | Onions | Salt | Spices", "Macaroni | Cheese Sauce | Milk | Butter | Salt | Pepper", "Ground Meat | Breadcrumbs | Onions | Garlic | Spices", "Chicken Tenders | Breading Mix | Vegetable Oil | Salt | Pepper" )), row.names = c(NA, 5L), class = "data.frame")
问题
需要找出Ingredients列中包含"Salt"的行,尝试执行:
df %>% str_detect(Ingredients, "Salt")
得到错误:Error: object 'Ingredients' not found,但将str_detect放入filter()中执行:
df %>% filter(str_detect(Ingredients, "Salt"))
却能正常返回匹配数据集。已知Ingredients是字符类型,为何直接传列名不行?放入filter后有何变化?
解答
str_detect()是stringr包的函数,它的第一个参数要求是字符向量,而非数据框。当你用df %>% str_detect(Ingredients, "Salt")时,管道符%>%会把整个df数据框传给str_detect()的第一个参数,而str_detect()无法处理数据框;同时,Ingredients是df的列,不是全局环境中的独立变量,所以函数找不到这个对象,报错。filter()是dplyr的核心数据操作函数,运行在tidy评估环境中。它会自动识别传入数据框中的列,把Ingredients作为字符向量传递给str_detect()的第一个参数,因此str_detect()能正常接收字符向量并执行匹配逻辑,最终filter()会根据匹配结果筛选出符合条件的行。
补充:如果想不用filter()直接得到匹配的布尔向量,可以用pull()先提取列:
df %>% pull(Ingredients) %>% str_detect("Salt")
内容的提问来源于stack exchange,提问作者Jay Bee
相关产品推荐
相关产品推荐

