基于正则表达式区分STRING与STRINGLIST数据类型的技术求助
Hey there, let’s work through your data type detection problem—both the regex tweaks you need and whether regex is the right tool for the job.
First: Is Regex a Good Fit Here?
Short answer: Yes, especially since your app already uses regex to map patterns to data types. Regex excels at structured text pattern matching, which aligns perfectly with your need to classify STRING/STRINGLIST based on delimiters and format. The only caveat is handling binary UTF16 data—you’ll need a quick preprocessing step to decode it before applying regex, which we’ll cover below.
Part 1: Regular Format Tweaks
Let’s fix and refine your regex for plaintext STRING/STRINGLIST:
STRING Detection
Your existing regex ^>\w+<$ works for alphanumeric-only STRINGs, but if your STRINGs can include non-alphanumeric characters (like spaces or symbols, as long as they don’t contain </>), update it to:
^>[^<>]+<$
This matches any string wrapped in >< that doesn’t contain nested < or >—exactly the STRING format you described.
STRINGLIST Detection
Your current regex catches <|> sequences but misses comma-separated cases. We can combine both conditions into a single regex that identifies STRINGLIST by either containing <|> OR having at least one comma (while ensuring it’s not a valid STRING):
^(?:[^<>]*<\|>[^<>]*|[^<>]+,[^<>]+)$
Breakdown:
[^<>]*<\|>[^<>]*: Matches any string with a<|>sequence (no nested</>)[^<>]+,[^<>]+: Matches any string with at least one comma (again, no</>to avoid conflicting with STRING)(?:...): Non-capturing group to combine the two conditions
Pro tip: Always run STRING detection first—if a string matches the STRING regex, skip checking for STRINGLIST to avoid false positives.
Part 2: Binary UTF16 Format Handling
UTF16 binary data needs decoding before regex can work (since regex operates on text, not raw bytes). Here’s how to approach it:
Step 1: Preprocess Binary Data
First, decode the binary bytes into a UTF16 string. You’ll need to know the byte order (big-endian or little-endian) of your data—most systems use little-endian, but confirm this first. For example, in Python you’d use bytes.decode('utf-16-le') or utf-16-be. If your data has a BOM (byte order mark), use utf-16 as the decoding format (it will auto-detect the byte order).
STRING Detection (UTF16)
Once decoded, use the same refined STRING regex from above (^>[^<>]+<$), but add a check that the content inside >< is readable UTF16 text (not control characters):
^>[\x20-\xD7FF\xE000-\xFFFD]+<$
This matches STRINGs with printable UTF16 characters (excludes unprintable control codes).
STRINGLIST Detection (UTF16)
After decoding, use a regex that detects comma-separated elements (since you mentioned STRINGLIST here is only identified by commas):
^[^<>]+(?:,[^<>]+)+$
This matches strings with one or more commas, no </> (to avoid confusion with STRING), and ensures there’s content on at least one side of the comma. If you need to allow empty elements (e.g., abc, or ,def), adjust it to ^[^<>]*(?:,[^<>]*)*$ (just make sure to exclude empty strings if needed).
Key Notes to Avoid Edge Cases
- Order of Operations: Always check for STRING first. If a string matches the STRING pattern, it can’t be a STRINGLIST—this prevents false positives.
- Nested Delimiters: If your data ever includes nested
</>characters, you’ll need a more complex regex (or a parser) to handle it, but based on your description, this shouldn’t be an issue.
内容的提问来源于stack exchange,提问作者user1808484

