R语言中如何正确标记二进制变量v4:忽略前置排除后置字符
Let's break down what's going wrong and how to fix it.
The Issue with Your Current Code
Your regex (a|b|c)$ only checks if the very last character of the string is a, b, or c. This works for simple cases like "a" or "2020c", but it fails to properly handle scenarios where a/b/c appears earlier in the string followed by other content (like "c 2020"). Based on your example expectations, what you really need is to ensure a/b/c is the final independent unit in the string—either the whole string, the last space-separated element, or the trailing character after prefixes like years.
The Fix
Update your regex to include a word boundary (\\b) to target a/b/c as a distinct, trailing unit:
df1$v4 <- as.integer(grepl("\\b(a|b|c)$", df1$v1))
How This Works
\\b: Matches a word boundary—this is the transition between a letter/number and a non-letter/number, or the start/end of the string. This ensures we're not matchinga/b/cas part of a longer word (like"c2020").(a|b|c)$: Ensures the match occurs exactly at the end of the string.
Testing this against your example cases:
"a"→ matches → 1"b"→ matches →1"d"→ no match →0"c 2020"→ ends with"2020"→ no match →0"2020 c"→ ends with standalone"c"→ matches →1"1990 c"→ ends with standalone"c"→ matches →1
This perfectly aligns with your desired output.
内容的提问来源于stack exchange,提问作者Max H.

