基于RegEx从特殊格式数据中提取用户名的技术问询
Looking at your data structure, fields are prefixed with sequences like 000 [digits] [field_code], and the username you need ("SARITEIO SARICO JR") sits right after the marker 000 212 118. Here's a targeted regex solution to pull it out:
Solution Regex
000 212 118([\w\s]+?)000
How It Works
Let’s break down each component to understand why this works:
000 212 118: This matches the exact prefix that precedes your username field. The118appears to be the consistent field identifier for the username entry.([\w\s]+?): This capturing group grabs the username itself:\wcovers alphanumeric characters (perfect for names)\saccounts for spaces between the first name, last name, and suffix+?makes the match non-greedy, so it stops immediately at the next field marker (000) instead of over-matching later content.
000: This is the start of the next field, acting as our clean stopping point.
Testing Against Your Sample Data
When run on your snippet:
...000 212 118SARITEIO SARICO JR000 212...
The regex will capture the group SARITEIO SARICO JR—exactly the username you’re targeting.
Flexible Variant
If the leading three-digit number (212) might vary but the field code (118) stays the same, use this more adaptable version:
000 \d{3} 118([\w\s]+?)000
Here, \d{3} matches any three-digit number, so it works even if the prefix digits change as long as the username field code remains 118.
This works with most modern regex engines (Python’s re, JavaScript’s RegExp, etc.) that support non-greedy quantifiers.
内容的提问来源于stack exchange,提问作者Sari Rahal

