You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Conduit逐行正则匹配及管道操作相关技术问题

Conduit 正则匹配与管道处理问题

根据示例,我们可以实现获取每行长度的功能:

import Conduit
import Data.Text (Text, pack)
import Text.Regex.TDFA ((=~), getAllTextMatches)
import Control.Monad.IO.Class (liftIO)

wc :: IO ()
wc = runResourceT
       $ runConduit
       $ sourceFile "input.txt"
       .| decodeUtf8C
       .| peekForeverE (lineC lengthCE >>= liftIO . print)

但我该如何基于正则获取所有匹配结果并最终写入文件?我尝试了以下代码,但不确定如何正确处理:

regex :: IO ()
regex = runResourceT
      $ runConduit
      $ sourceFile "input.txt"
      .| decodeUtf8C
      .| do
         line <- mapCE (\l -> getAllTextMatches (l =~ "^foo") :: [Text])
         liftIO $ print $ line

更新1:打印行的同时不消耗,让其继续在管道传递

我发现了内置的lines函数,但有没有办法打印一行的同时不消耗它,让它继续在管道中传递?

我写了这段代码,它会逐行打印,但stdoutC最终没有输出内容:

grep :: IO ()
grep = runResourceT
    $ runConduit
    $ yield "foo\ndoo"
    .| decodeUtf8C
    .| Data.Conduit.Text.lines
    .| mapC (\a -> a =~ ("[fd]oo" :: Text))
    .| mapM_C (liftIO . (print :: Text -> IO ()))
    .| encodeUtf8C
    .| stdoutC

运行结果:

ghci> grep
"foo"
"doo"

更新2:解决打印问题,但疑惑await的顺序

我已经解决了管道中打印的问题,代码如下:

grep :: IO ()
grep = runResourceT
    $ runConduit
    $ yieldMany ["foo\ndoo", "\nduh"]
    .| decodeUtf8C
    .| Data.Conduit.Text.lines
    .| mapC (\a -> a =~ ("[fd]oo" :: Text) :: Text)
    .| log1
    .| unlinesC
    .| encodeUtf8C
    .| stdoutC

但我有个疑问:为什么await的顺序很重要?比如下面这个log1函数里,await必须放在第一位?

log1 :: ConduitT Text Text (ResourceT IO) ()
log1 = do
       Just l <- await -- <- 必须放在第一位
       liftIO $ print l
       yield l

内容的提问来源于stack exchange,提问作者phoxd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 09:05:23