含特殊字符的Cortex型号正则匹配失效,请求优化正则表达式
Hey there, I see the problem—your current regex fails to capture the full Cortex-AXX pattern when trademark symbols (like ® or ™) or other special characters sit between Cortex and the model number. Let's tweak the regex to handle those pesky interfering characters properly.
Improved Regex Pattern
Here's a revised regex that skips over any non-alphanumeric, non-hyphen characters between the parts of the model name:
(?i)(Cortex)\b[^\w-]*-?[^\w-]*(\bA\b)[^\w-]*-?[^\w-]*(\d{1,2})
If you want something a bit more concise (and still effective):
(?i)Cortex\s*[^\w]*-?[^\w]*A[^\w]*-?[^\w]*(\d{1,2})
Breakdown of the Regex
Let's walk through each part so you understand how it works:
(?i): Turns on case-insensitive matching, so it catchescortex,Cortex,CORTEX—no matter the capitalization.(Cortex)\b: Captures the core name, with a word boundary (\b) to avoid accidental matches likeCortexX.[^\w-]*: Matches any number of characters that aren't letters, numbers, or hyphens—this covers those®,™, random spaces, or other special symbols messing up your match.-?: Matches an optional hyphen, since some models are written asCortex A53and others asCortex-A53.(\bA\b): Captures theAprefix with word boundaries to ensure we're not matching part of another word.(\d{1,2}): Grabs the 1-2 digit model number (like 53, 54, etc.).
Testing It Out
Let's test this against your example inputs:
- Clean input:
Helio P10 processor ARM cortex-A53 1.8GHz- Match result:
cortex-A53(with capture groups forcortex,A, and53)
- Match result:
- Input with special chars:
Helio P10 processor ARM® cortex®-A53™1.8GHz- Match result:
cortex®-A53(if you want to strip the special symbols afterward, you can use a simple string replace, or adjust the regex to only capture the core parts and reconstruct the clean format)
- Match result:
If you want to directly get a clean Cortex-AXX string without any special symbols, use this regex and combine the capture groups:
(?i)(Cortex)\b.*?(\bA\b).*?(\d{1,2})
Then just concatenate the groups with a hyphen: $1-$2$3 will give you Cortex-A53 even from the messy input.
Quick Tip
If you run into other special characters (like underscores or periods), you can adjust the [^\w-] part to [^A-Za-z0-9-] to cover all bases. For broader Unicode support (if dealing with non-Latin special chars), use \P{L} instead to match any non-letter characters—it works better in regex engines that support Unicode properties.
内容的提问来源于stack exchange,提问作者Bhagwati Malav

