使用RegEx提取零件规格遇特殊字符问题求助
Hey there! I see exactly what's tripping up your regex here. The problem is that your original customer data uses Cyrillic characters instead of Latin ones for the "M" and "x":
- The "М" in
М4х20is a Cyrillic capital M (Unicode U+041C), not the standard Latin M (U+004D) - The "х" is a Cyrillic small x (Unicode U+0445), not the Latin x (U+0078)
Your existing regex only looks for the Latin versions, so it naturally misses these Cyrillic variants in the global data.
Solution: Update Your Regex to Cover Both Character Sets
We can adjust the regex to include both Latin and Cyrillic equivalents using character classes. Here's how to do it:
Example Modified Regex
If your original regex looked something like this (to capture the M\d+x\d+ pattern):
(M\d+x\d+)
Update it to this expanded version:
([MmМм]\d+[XxХх]\d+)
Breakdown of the changes:
[MmМм]: Matches Latin uppercase M, Latin lowercase m, Cyrillic uppercase М, or Cyrillic lowercase м[XxХх]: Matches Latin uppercase X, Latin lowercase x, Cyrillic uppercase Х, or Cyrillic lowercase х
Bonus: Normalize to Latin Characters (If Needed)
If you want to standardize all extracted values to use Latin characters (so М4х20 becomes M4x20 for consistency), add a simple replacement step after extraction. For example, in Python:
import re text = "Screw М4х20 , DIN7985" match = re.search(r'([MmМм]\d+[XxХх]\d+)', text) if match: normalized_spec = match.group(1).replace('М', 'M').replace('м', 'm').replace('Х', 'X').replace('х', 'x') print(normalized_spec) # Outputs: M4x20
Testing It Out
This updated regex will now correctly extract the spec from both versions:
- For
Screw M4x20 , DIN7985→ capturesM4x20 - For
Screw М4х20 , DIN7985→ capturesМ4х20(or normalized toM4x20if you use the replacement step)
内容的提问来源于stack exchange,提问作者Martin Viczian

