Haskell中多次运行RegModule Monad并更新状态的问题
问题详情
尝试在Haskell中使用RegModule Monad,让前一次的状态影响下一次执行,但发现每次调用runRegModule时,位置增量仅首次生效,后续未累加。希望解决状态正确更新的问题,同时想知道是否有更简便的方式(比如用forM)让Monad自动运行至失败。
代码实现
import qualified Data.Set as S import Control.Monad type CharSet = S.Set Char data RE = RClass Bool CharSet newtype RegModule d a = RegModule {runRegModule :: String -> Int -> d -> [(a, Int, d)]} instance Monad (RegModule d) where return a = RegModule (\_s _i d -> return (a, 0, d)) m >>= f = RegModule (\s i d -> do (a, j, d') <- runRegModule m s i d (b, j', d'') <- runRegModule (f a) s (i + j) d' return (b, j + j', d'')) instance Functor (RegModule d) where fmap = liftM instance Applicative (RegModule d) where pure = return; (<*>) = ap scanChar :: RegModule d Char scanChar = RegModule (\s i d -> case drop i s of (c:cs) -> return (c, 1, d) [] -> [] ) regfail :: RegModule d a regfail = RegModule (\_s _i d -> [] ) regEX :: RE -> RegModule [String] () regEX (RClass b cs) = do next <- scanChar if (S.member next cs) then return () else regfail runRegModuleThrice :: RegModule d a -> String -> Int -> d -> [(a, Int, d)] runRegModuleThrice matcher input startPos state = let (result1, pos1, newState1) = head $ runRegModule matcher input startPos state (result2, pos2, newState2) = head $ runRegModule matcher input pos1 newState1 (result3, pos3, newState3) = head $ runRegModule matcher input pos2 newState2 in [(result1, pos1, newState1), (result2, pos2, newState2), (result3, pos3, newState3)]
实际输出
ghci> runRegModuleThrice (regEX (RClass False (S.singleton 'a'))) "aaa" 0 [] [((),1,[]),((),1,[]),((),1,[])]
预期输出
ghci> runRegModuleThrice (regEX (RClass False (S.singleton 'a'))) "aaa" 0 [] [((),1,[]),((),2,[]),((),3,[])]
补充疑问
- 若因缺少幺半群值导致状态无法积累,是否会丢失Monad核心特性(如《Learn You a Haskell》中Writer Monad特性)?
- 不修改Monad定义,能否展示积累的状态或追加文本?
- 是否仅修改
regEX、scanChar、regfail或runRegModuleThrice就能实现类似[((),1,['a']),((),2,['a']),((),3,['a'])]的输出? - 为何《Learn You a Haskell》示例未强调修改Monad运行器?是否与Haskell版本有关?
解决方案
核心问题分析
你的runRegModuleThrice里,每次提取的pos1、pos2是当前匹配的增量(比如每次匹配一个字符返回1),而不是累计的绝对位置。RegModule的绑定逻辑中,runRegModule返回的第二个值是本次操作的位置偏移量,不是绝对位置。因此需要手动累加绝对位置,而不是直接把偏移量作为下一次的起始位置。
修改runRegModuleThrice实现
把每次的起始位置和偏移量累加得到下一次的起始位置,同时记录累计后的绝对位置:
runRegModuleThrice :: RegModule d a -> String -> Int -> d -> [(a, Int, d)] runRegModuleThrice matcher input startPos state = let (result1, offset1, newState1) = head $ runRegModule matcher input startPos state absPos1 = startPos + offset1 (result2, offset2, newState2) = head $ runRegModule matcher input absPos1 newState1 absPos2 = absPos1 + offset2 (result3, offset3, newState3) = head $ runRegModule matcher input absPos2 newState2 absPos3 = absPos2 + offset3 in [(result1, absPos1, newState1), (result2, absPos2, newState2), (result3, absPos3, newState3)]
运行后就能得到预期输出,因为每次传递的是累计后的绝对位置,返回的也是累计后的位置值。
自动运行至失败的实现(用forM)
可以通过递归结合forM实现循环执行,直到匹配失败:
import Control.Monad (forM) runUntilFail :: RegModule d a -> String -> Int -> d -> [(a, Int, d)] runUntilFail matcher input startPos state = let step (pos, s) = case runRegModule matcher input pos s of [] -> Nothing [(res, off, s')] -> Just ((res, pos + off, s'), (pos + off, s')) steps = iterate (step . snd) (step (startPos, state)) in takeWhile isJust $ map fst $ catMaybes steps
也可以用更简洁的递归实现,避免处理无限序列的问题:
runUntilFail :: RegModule d a -> String -> Int -> d -> [(a, Int, d)] runUntilFail matcher input startPos state = case runRegModule matcher input startPos state of [] -> [] [(res, offset, newState)] -> let absPos = startPos + offset in (res, absPos, newState) : runUntilFail matcher input absPos newState
补充疑问解答
关于幺半群与Monad特性:
不会丢失Monad的核心特性。Monad的核心是满足绑定律、左单位元、右单位元的规则,和是否携带幺半群状态无关。Writer Monad是带输出的Monad,需要幺半群来合并输出,但你的RegModule是带位置偏移和自定义状态的Monad,状态更新逻辑由你自行定义,不需要依赖幺半群。不修改Monad定义实现状态积累:
可以。将状态d设计为累计型(比如[Char]),在scanChar或regEX中更新状态:scanChar :: RegModule [Char] Char scanChar = RegModule (\s i d -> case drop i s of (c:cs) -> return (c, 1, c : d) [] -> [] )每次匹配后状态会累计字符,运行后就能得到类似
[((),1,['a']),((),2,['a','a']),((),3,['a','a','a'])]的输出(若需要正序可以用d ++ [c],但效率较低)。仅修改指定函数实现目标输出:
可以。修改scanChar更新状态,同时配合修改后的runRegModuleThrice累计绝对位置:scanChar :: RegModule [String] Char scanChar = RegModule (\s i d -> case drop i s of (c:cs) -> return (c, 1, ["a"] ++ d) [] -> [] )运行后就能得到你想要的输出格式。
《Learn You a Haskell》示例的差异:
和Haskell版本无关。LYAH中的示例(如State、Writer Monad)通常提供封装好的运行器(比如runState、runWriter),这些运行器已经处理了状态累计逻辑。而你的RegModule是自定义Monad,runRegModule返回的是本次偏移和新状态,需要手动串联累加,这是因为该Monad的设计和标准库State Monad不同——State Monad的运行器返回最终状态,而RegModule返回的是单次操作的偏移量。
内容的提问来源于stack exchange,提问作者Piskator

