do符号从Monad解包的值及决定因素,RegModule实例困惑解析
next <- scanChar的绑定逻辑 问题描述
刚接触Monad不久,发现此前的理解存在偏差。在以下最小可复现代码中,难以理解next <- scanChar为何能返回单个Char并绑定到next。按之前的认知,do符号会解包Monad内的所有值,绑定到next的应该是[(Char, Int, d)]才对,是否是构造函数或实例定义导致了这种差异?
import qualified Data.Set as S import Control.Monad type CharSet = S.Set Char data RE = RClass Bool CharSet newtype RegModule d a = RegModule {runRegModule :: String -> Int -> d -> [(a, Int, d)]} instance Monad (RegModule d) where return a = RegModule (\_s _i d -> return (a, 0, d)) m >>= f = RegModule (\s i d -> do (a, j, d') <- runRegModule m s i d (b, j', d'') <- runRegModule (f a) s (i + j) d' return (b, j + j', d'')) instance Functor (RegModule d) where fmap = liftM instance Applicative (RegModule d) where pure = return; (<*>) = ap scanChar :: RegModule d Char scanChar = RegModule (\s i d -> case drop i s of (c:cs) -> return (c, 1, d) [] -> [] ) regfail :: RegModule d a regfail = RegModule (\_s _i d -> [] ) regEX :: RE -> RegModule [String] () regEX (RClass b cs) = do next <- scanChar if (S.member next cs) then return () else regfail runRegModuleThrice :: RegModule d a -> String -> Int -> d -> [(a, Int, d)] runRegModuleThrice matcher input startPos state = let (result1, pos1, newState1) = head $ runRegModule matcher input startPos state (result2, pos2, newState2) = head $ runRegModule matcher input pos1 newState1 (result3, pos3, newState3) = head $ runRegModule matcher input pos2 newState2 in [(result1, pos1, newState1), (result2, pos2, newState2), (result3, pos3, newState3)]
解答
你混淆了RegModule这个Monad本身和它内部包含的列表结构。next <- scanChar能绑定到单个Char,完全是由RegModule d的Monad实例定义决定的,核心是>>=的实现逻辑。
关键逻辑拆解
RegModule的本质:它是一个封装了「带状态、输入位置的非确定性计算」的Monad。
runRegModule返回的[(a, Int, d)]代表该计算可能产生多个结果,每个结果包含三个部分:计算出的值a、消耗的输入长度Int、更新后的状态d。Monad实例的
>>=做了什么:
当你在do notation中写next <- scanChar,本质是调用scanChar >>= \next -> ...。看>>=的具体实现:m >>= f = RegModule (\s i d -> do (a, j, d') <- runRegModule m s i d -- 这里是List Monad的do,遍历m的所有结果元组 (b, j', d'') <- runRegModule (f a) s (i + j) d' -- 用元组中的a生成新计算,遍历其结果 return (b, j + j', d'') )内层的
do是List Monad的语法,负责遍历runRegModule m返回的所有结果元组,把元组的第一个元素a(也就是scanChar产生的Char)传给后续的函数f(即你写的\next -> ...逻辑)。为什么绑定的是单个Char而非列表:
RegModule的Monad实例已经把内部的列表遍历逻辑完全封装了。在RegModule的do notation中,<-绑定的是单个计算结果的值(即元组中的a),而非整个结果列表。列表在这里代表计算的非确定性——如果一个RegModule计算返回多个结果,Monad会自动处理所有分支的后续逻辑,最终合并所有可能的结果。
总结
你之前的误解是把Monad内部的底层结构(列表)当成了Monad要暴露给用户的值,但实际上Monad的>>=定义决定了用户层能拿到什么。RegModule d的Monad实例帮你处理了非确定性的列表遍历,所以在do块中绑定到的是单个计算值,而非整个结果列表。
内容的提问来源于stack exchange,提问作者Piskator

