R语言正则需求:移除单词末尾的s,但排除前导为i或s的情况
Hey there! I see you've been working with R for a while and are looking to level up your regex skills—this is a perfect practical problem to do just that. Your current sub("s$","",dfx) works for stripping any trailing 's', but we need to add a condition to skip cases where the 's' comes right after an 'i' or 's'.
The Fix: Use a Negative Lookbehind Assertion
We can leverage a negative lookbehind in regex to tell R: "only remove the trailing 's' if it's NOT immediately preceded by 'i' or 's'". Here's the adjusted code:
dfx <- c("class","guests","trips","sassy","basis","fruits") dfx_fixed <- sub("(?<![is])s$", "", dfx, perl = TRUE) dfx_fixed
Let's Break Down the Regex
(?<![is]): This is the negative lookbehind. It checks that the character right before the trailing 's' is not 'i' or 's'. The[is]is a character class that matches either of those two letters.s$: Matches a literal 's' at the very end of the string (the$denotes the end of the string).perl = TRUE: R's base regex functions require this flag to enable lookbehind assertions, which aren't supported in basic regex mode.
Test the Result
When you run the code above, you'll get this output:
[1] "class" "guest" "trip" "sassy" "basis" "fruit"
Let's verify each case matches your requirements:
- "class": Trailing 's' is preceded by 's' → remains unchanged
- "guests": Trailing 's' is preceded by 't' → stripped to "guest"
- "trips": Trailing 's' is preceded by 'p' → stripped to "trip"
- "sassy": Doesn't end with 's' → remains unchanged
- "basis": Trailing 's' is preceded by 'i' → remains unchanged
- "fruits": Trailing 's' is preceded by 't' → stripped to "fruit"
Bonus: Handling Multiple Trailing 's'
If you ever need to strip all trailing 's' (not just the final one) except when the preceding character is 'i' or 's', you can adjust the regex to (?<![is])s+$—but this is probably beyond your current need.
Hope this helps you nail this problem and build your regex confidence!
内容的提问来源于stack exchange,提问作者user1185787

