如何用Conduit逐行正则匹配及管道操作相关技术问题
Conduit 正则匹配与管道处理问题
根据示例,我们可以实现获取每行长度的功能:
import Conduit import Data.Text (Text, pack) import Text.Regex.TDFA ((=~), getAllTextMatches) import Control.Monad.IO.Class (liftIO) wc :: IO () wc = runResourceT $ runConduit $ sourceFile "input.txt" .| decodeUtf8C .| peekForeverE (lineC lengthCE >>= liftIO . print)
但我该如何基于正则获取所有匹配结果并最终写入文件?我尝试了以下代码,但不确定如何正确处理:
regex :: IO () regex = runResourceT $ runConduit $ sourceFile "input.txt" .| decodeUtf8C .| do line <- mapCE (\l -> getAllTextMatches (l =~ "^foo") :: [Text]) liftIO $ print $ line
更新1:打印行的同时不消耗,让其继续在管道传递
我发现了内置的lines函数,但有没有办法打印一行的同时不消耗它,让它继续在管道中传递?
我写了这段代码,它会逐行打印,但stdoutC最终没有输出内容:
grep :: IO () grep = runResourceT $ runConduit $ yield "foo\ndoo" .| decodeUtf8C .| Data.Conduit.Text.lines .| mapC (\a -> a =~ ("[fd]oo" :: Text)) .| mapM_C (liftIO . (print :: Text -> IO ())) .| encodeUtf8C .| stdoutC
运行结果:
ghci> grep "foo" "doo"
更新2:解决打印问题,但疑惑await的顺序
我已经解决了管道中打印的问题,代码如下:
grep :: IO () grep = runResourceT $ runConduit $ yieldMany ["foo\ndoo", "\nduh"] .| decodeUtf8C .| Data.Conduit.Text.lines .| mapC (\a -> a =~ ("[fd]oo" :: Text) :: Text) .| log1 .| unlinesC .| encodeUtf8C .| stdoutC
但我有个疑问:为什么await的顺序很重要?比如下面这个log1函数里,await必须放在第一位?
log1 :: ConduitT Text Text (ResourceT IO) () log1 = do Just l <- await -- <- 必须放在第一位 liftIO $ print l yield l
内容的提问来源于stack exchange,提问作者phoxd
相关产品推荐
相关产品推荐

