You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Kedro Catalog中添加并批量读取多个.md文件

Kedro批量管理并读取多份.md文件

要实现把指定目录下的所有.md文件在Kedro Catalog里用单个名称统一管理,你需要自定义数据集扩展Kedro的功能,具体步骤如下:

1. 自定义Markdown批量读取数据集

在你的Kedro项目的src/<project_name>/datasets/目录下新建Python文件(比如markdown_batch.py),编写自定义数据集类:

from pathlib import Path
from typing import Dict
from kedro.io import AbstractDataSet

class MarkdownBatchDataSet(AbstractDataSet):
    def __init__(self, filepath: str):
        self._filepath = Path(filepath)

    def _load(self) -> Dict[str, str]:
        # 遍历目录下所有.md文件并读取内容
        md_files = list(self._filepath.glob("*.md"))
        md_content = {}
        for file in md_files:
            with open(file, "r", encoding="utf-8") as f:
                md_content[file.name] = f.read()
        return md_content

    def _save(self, data: Dict[str, str]) -> None:
        # 若不需要保存功能,直接抛出未实现异常
        raise NotImplementedError("MarkdownBatchDataSet不支持保存操作")

    def _describe(self) -> Dict[str, str]:
        return {"filepath": str(self._filepath)}

2. 在Catalog中配置自定义数据集

打开conf/base/catalog.yml,添加如下配置,将data/01_raw/folder_name/路径下的所有.md文件统一用raw_markdown_files名称管理:

raw_markdown_files:
  type: <project_name>.datasets.markdown_batch.MarkdownBatchDataSet
  filepath: data/01_raw/folder_name/

注意把<project_name>替换成你实际的Kedro项目名称。

3. 在Pipeline中调用数据集

你可以在Pipeline节点里直接使用这个数据集,示例代码如下:

from kedro.pipeline import Pipeline, node
from .nodes import process_markdown_files

def create_pipeline(**kwargs) -> Pipeline:
    return Pipeline(
        [
            node(
                func=process_markdown_files,
                inputs="raw_markdown_files",
                outputs="processed_markdown",
                name="process_markdown_node",
            )
        ]
    )

对应的节点处理函数示例:

def process_markdown_files(md_files: Dict[str, str]) -> Dict[str, dict]:
    processed = {}
    for filename, content in md_files.items():
        # 这里写你的业务处理逻辑,比如提取标题、解析内容
        lines = content.split("\n")
        title = next((line.strip("# ") for line in lines if line.startswith("#")), filename)
        processed[filename] = {"title": title, "content": content}
    return processed

这样就能通过raw_markdown_files这个Catalog名称,一次性获取指定目录下所有.md文件的内容了。

内容的提问来源于stack exchange,提问作者Lucas Jurani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 00:18:11