You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于正则表达式区分STRING与STRINGLIST数据类型的技术求助

Regex Solutions for Distinguishing STRING vs STRINGLIST (Plus Feasibility Check)

Hey there, let’s work through your data type detection problem—both the regex tweaks you need and whether regex is the right tool for the job.

First: Is Regex a Good Fit Here?

Short answer: Yes, especially since your app already uses regex to map patterns to data types. Regex excels at structured text pattern matching, which aligns perfectly with your need to classify STRING/STRINGLIST based on delimiters and format. The only caveat is handling binary UTF16 data—you’ll need a quick preprocessing step to decode it before applying regex, which we’ll cover below.


Part 1: Regular Format Tweaks

Let’s fix and refine your regex for plaintext STRING/STRINGLIST:

STRING Detection

Your existing regex ^>\w+<$ works for alphanumeric-only STRINGs, but if your STRINGs can include non-alphanumeric characters (like spaces or symbols, as long as they don’t contain </>), update it to:

^>[^<>]+<$

This matches any string wrapped in >< that doesn’t contain nested < or >—exactly the STRING format you described.

STRINGLIST Detection

Your current regex catches <|> sequences but misses comma-separated cases. We can combine both conditions into a single regex that identifies STRINGLIST by either containing <|> OR having at least one comma (while ensuring it’s not a valid STRING):

^(?:[^<>]*<\|>[^<>]*|[^<>]+,[^<>]+)$

Breakdown:

  • [^<>]*<\|>[^<>]*: Matches any string with a <|> sequence (no nested </>)
  • [^<>]+,[^<>]+: Matches any string with at least one comma (again, no </> to avoid conflicting with STRING)
  • (?:...): Non-capturing group to combine the two conditions

Pro tip: Always run STRING detection first—if a string matches the STRING regex, skip checking for STRINGLIST to avoid false positives.


Part 2: Binary UTF16 Format Handling

UTF16 binary data needs decoding before regex can work (since regex operates on text, not raw bytes). Here’s how to approach it:

Step 1: Preprocess Binary Data

First, decode the binary bytes into a UTF16 string. You’ll need to know the byte order (big-endian or little-endian) of your data—most systems use little-endian, but confirm this first. For example, in Python you’d use bytes.decode('utf-16-le') or utf-16-be. If your data has a BOM (byte order mark), use utf-16 as the decoding format (it will auto-detect the byte order).

STRING Detection (UTF16)

Once decoded, use the same refined STRING regex from above (^>[^<>]+<$), but add a check that the content inside >< is readable UTF16 text (not control characters):

^>[\x20-\xD7FF\xE000-\xFFFD]+<$

This matches STRINGs with printable UTF16 characters (excludes unprintable control codes).

STRINGLIST Detection (UTF16)

After decoding, use a regex that detects comma-separated elements (since you mentioned STRINGLIST here is only identified by commas):

^[^<>]+(?:,[^<>]+)+$

This matches strings with one or more commas, no </> (to avoid confusion with STRING), and ensures there’s content on at least one side of the comma. If you need to allow empty elements (e.g., abc, or ,def), adjust it to ^[^<>]*(?:,[^<>]*)*$ (just make sure to exclude empty strings if needed).


Key Notes to Avoid Edge Cases

  • Order of Operations: Always check for STRING first. If a string matches the STRING pattern, it can’t be a STRINGLIST—this prevents false positives.
  • Nested Delimiters: If your data ever includes nested </> characters, you’ll need a more complex regex (or a parser) to handle it, but based on your description, this shouldn’t be an issue.

内容的提问来源于stack exchange,提问作者user1808484

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:33:44