如何用Parser Combinators实现无歧义的字符串修剪?
解决ReadP实现trim函数的歧义问题
我用Text.ParserCombinators.ReadP实现的trim函数如下:
trim :: ReadP String trim = do skipSpaces content <- some get skipSpaces eof return content
运行时总会出现歧义结果:
ghci> readP_to_S trim " hello world " [("hello world",""),("hello world ","")]
我知道问题出在需要保证修剪后的内容不以空格结尾,但目前只想到一种效率极低的实现方式:
trim :: ReadP String trim = do skipSpaces content <- many get lastOne <- satisfy (not . isSpace) skipSpaces eof return (content ++ [lastOne])
有没有更好的实现方案?
方案一:用suchThat过滤解析结果
利用suchThat对解析出的结果进行过滤,确保内容非空且最后一个字符不是空格,同时保留中间的空格:
import Text.ParserCombinators.ReadP import Data.Char (isSpace) trim :: ReadP String trim = do skipSpaces content <- some get `suchThat` (\s -> not (null s) && not (isSpace (last s))) skipSpaces eof return content
测试结果:
ghci> readP_to_S trim " hello world " [("hello world","")]
这种方式代码简洁,逻辑清晰,能直接消除歧义。
方案二:精确解析避免回溯
如果追求更高效率,可以直接解析出符合要求的内容,避免事后过滤的开销:
import Text.ParserCombinators.ReadP import Data.Char (isSpace) trim :: ReadP String trim = do skipSpaces content <- singleNonSpace <|> multiCharContent skipSpaces eof return content where singleNonSpace = (:[]) <$> satisfy (not . isSpace) multiCharContent = do start <- satisfy (not . isSpace) rest <- manyTill get (satisfy (not . isSpace)) return (start : rest)
这个实现通过分支处理单字符和多字符场景,直接解析出首尾非空格、中间任意的内容,完全避免了回溯和过滤,效率更高。
内容的提问来源于stack exchange,提问作者Futarimiti
相关产品推荐
相关产品推荐

