基于regex.h的C程序正则校验需求咨询及现有方案确认
regex.h Hey there! Let's break down your two regex validation needs and get them sorted properly for your C code.
1. String must start and end with @, only two @ total (none in the middle), with alphanumerics in between
Your current regex ^[@]+[A-Za-z0-9]+[@]$ has a small issue: the [@]+ at the start will match one or more @ characters. That means strings like @@abc@ would incorrectly pass the check, violating your "only two @ total" rule.
Correct Regex:
regexCheck(tag, "^@[A-Za-z0-9]+@$") == 1
Breakdown:
^@: Anchors to the start of the string and requires exactly one @ (no more, no less)[A-Za-z0-9]+: Matches one or more letters/numbers—since @ isn't in this character set, it ensures no @ characters appear in the middle@$: Anchors to the end of the string and requires exactly one @ to close things out
This guarantees exactly two @ characters (one at start, one at end) with only alphanumerics in between.
2. String must contain at least one underscore, rest are alphanumerics
Your current regex [A-Za-z0-9_] is way too loose. It only checks if any single character in the string is alphanumeric or an underscore—so even a string like test (no underscores) would pass, because it has matching characters. We need to validate the entire string and enforce the underscore requirement.
Correct Regex (depends on your regex flavor):
Most C regex.h usage uses POSIX regexes, so we'll cover both common cases:
If using POSIX Extended Regular Expressions (ERE, enabled with REG_EXTENDED flag in regcomp):
regexCheck(string, "^(?=.*_)[A-Za-z0-9_]+$") == 1
Breakdown:
^and$: Anchor to the start and end of the string, ensuring we validate the entire input (not just a substring)(?=.*_): A positive lookahead that checks the string contains at least one underscore (.*matches any number of any characters, followed by_)[A-Za-z0-9_]+: Matches one or more alphanumeric characters or underscores, ensuring no invalid characters are present
If using POSIX Basic Regular Expressions (BRE, default if no REG_EXTENDED):
BRE doesn't support lookaheads, so we use a pattern that explicitly requires an underscore somewhere:
regexCheck(string, "^[A-Za-z0-9]*_[A-Za-z0-9]*$") == 1
Breakdown:
^and$: Again, anchor to the full string[A-Za-z0-9]*: Matches zero or more alphanumerics (allows the underscore to be at the start)_: Requires exactly one underscore (the*around it means we can have alphanumerics before, after, or both)[A-Za-z0-9]*: Matches zero or more alphanumerics after the underscore (allows the underscore to be at the end)
This works for all valid cases: _test, te_st, test_ will all match, while test (no underscore) or test! (invalid character) will fail.
内容的提问来源于stack exchange,提问作者Andre D'Souza

