Haskell初学者求教:TDFA中标记转换的标记操作实现子正则匹配
Great job getting started with Haskell and TDFA for regex matching! Let's walk through how to use TDFA's tag system to capture subexpressions—this will make extracting specific parts of your matches way cleaner than relying on positional indices.
First, let's clarify the correct syntax for tagging subexpressions in TDFA: you wrap the subexpression you want to capture with (?<tagname>...) (alternatively (?'tagname'...) works too). Your original regex comment had (tag1 b|tag2 bb|...)—that's not the right tag syntax, so we'll adjust that first to properly mark each b/bb/etc. as tagged subexpressions.
Here's a step-by-step example with code:
Step 1: Import Required Modules
You'll need the core TDFA module plus a helper module for working with tags:
import Text.Regex.TDFA import Text.Regex.TDFA.Tags (getTag)
Step 2: Write a Tagged Regex
Let's rewrite your regex to tag each subexpression properly. We'll wrap each b/bb/bbb/bbbb in a named tag while keeping your original repetition behavior:
-- Tagged regex: each target subexpression gets a named tag taggedRegex :: String taggedRegex = "((?<tag1>b)|(?<tag2>bb)|(?<tag3>bbb)|(?<tag4>bbbb))*"
Step 3: Match and Extract Tagged Content
Use matchAllText to get all matches (including tag metadata), then use getTag to pull out the content for each tag by name:
main :: IO () main = do let targetString = "bb bbbb b bb" -- Compile the regex (type annotation ensures correct type inference) compiledRegex = makeRegex taggedRegex :: Regex -- Get all matches with tag information allMatches = matchAllText compiledRegex targetString -- Iterate through each match and print detailed results mapM_ printMatchDetails allMatches where printMatchDetails :: MatchText String -> IO () printMatchDetails match = do -- The first element of MatchText is the full matched substring let fullMatch = fst (head match) putStrLn $ "\nFull matched segment: " ++ fullMatch -- Check each tag and print its content if it was matched mapM_ printTagContent ["tag1", "tag2", "tag3", "tag4"] where printTagContent tag = case getTag tag match of Just (_, matchedText) -> putStrLn $ " " ++ tag ++ ": " ++ matchedText Nothing -> putStrLn $ " " ++ tag ++ ": No match in this segment"
Key Notes:
MatchText Stringholds the full match plus all tagged (and untagged) subexpression matches. ThegetTagfunction lets you directly access a tag's content by name—no need to count group positions, which is a huge win for readability.- If a tag doesn't match in a given segment (e.g.,
tag3in a segment that only matchedbb),getTagreturnsNothing—we handle that case gracefully in the example. - Tags are case-sensitive, so make sure your tag names in
getTagexactly match what's defined in the regex.
When you run this code, you'll see output like:
Full matched segment: bb tag1: No match in this segment tag2: bb tag3: No match in this segment tag4: No match in this segment Full matched segment: bbbb tag1: No match in this segment tag2: No match in this segment tag3: No match in this segment tag4: bbbb Full matched segment: b tag1: b tag2: No match in this segment tag3: No match in this segment tag4: No match in this segment Full matched segment: bb tag1: No match in this segment tag2: bb tag3: No match in this segment tag4: No match in this segment
That's it! This approach makes your subexpression extraction much more maintainable, especially as your regexes grow in complexity.
内容的提问来源于stack exchange,提问作者imran7

