使用stringr包的str_detect函数:精准匹配单词shop并排除shopping
Let's break down how to solve this string matching problem with R's stringr and dplyr packages.
The Core Problem
The default str_detect(..., "shop") matches any occurrence of the substring "shop"—including those inside "shopping". We need to target only the standalone word "shop", regardless of whether "shopping" is present.
The Right Regular Expression
We'll use word boundaries (\\b) to ensure we match "shop" as a complete word. This automatically excludes "shopping" because "shop" in "shopping" isn't followed by a word boundary (since "p" is a word character). The regex is:
"\\bshop\\b"
\\bmarks the position between a word character (like letters/numbers) and a non-word character (spaces, periods, exclamation points, etc.).- This regex only matches the standalone word "shop", ignoring partial matches like "shopping" or "shopkeeper".
Full Code Implementation
Here's how to integrate this into your workflow, including converting the remarks column to lowercase (to ensure case-insensitive matching, which you mentioned already doing):
library(dplyr) library(stringr) # Load your provided example data example <- structure(list(price = c(195000, 213000, 215000, 240000, 241000, 250000, 255000, 256500, 260000, 263500, 265000, 277000, 280000, 280000, 150000), remarks = c("large home with a 1200 sf shop. great location close to shopping.", "updated home close to shopping & schools.", "nice location. 2br home with updating.", "huge shop on property!", "close to shopping.", "updated, clean, great location, garage.", "close to shopping and massive shop on property.", "updated home near shopping, schools, restaurants.", "large home with updated interior.", "close to schools, updated, stick-built shop 1500sf.", "home and shop.", "near schools, shopping, restaurants. partially updated home.", "located close to shopping. high quality home with shop in backyard.", "brick 2-story. lots of shopping near by. detached garage and large shop in backyard.", "fixer! needs work.")), row.names = c(NA, -15L), class = c("tbl_df", "tbl", "data.frame")) # Create the shop_YN column example_with_shop <- example %>% mutate( remarks_lower = str_to_lower(remarks), # Included for completeness since you mentioned this step shop_YN = if_else(str_detect(remarks_lower, "\\bshop\\b"), "Yes", "No") ) # Check the results print(example_with_shop %>% select(remarks, shop_YN))
Testing Against Your Rules
Let's confirm this code meets all your requirements:
- Only "shop" present: Returns "Yes" (e.g., row 4: "huge shop on property!")
- Only "shopping" present: Returns "No" (e.g., row 2: "updated home close to shopping & schools.")
- Neither present: Returns "No" (e.g., row 3: "nice location. 2br home with updating.")
- Both "shop" and "shopping" present: Returns "Yes" (e.g., row 1: "large home with a 1200 sf shop. great location close to shopping.")
Edge Case Alternative Regex
If you need to explicitly rule out "shopping" even in edge cases (like unusual formatting), you can add a negative lookahead to ensure "shop" isn't followed by "ping":
"\\bshop(?!ping)\\b"
This adds an extra layer of specificity, though the simpler \\bshop\\b works perfectly for your example data and stated rules.
内容的提问来源于stack exchange,提问作者EastBeast

