如何在Kedro Catalog中添加并批量读取多个.md文件
Kedro批量管理并读取多份.md文件
要实现把指定目录下的所有.md文件在Kedro Catalog里用单个名称统一管理,你需要自定义数据集扩展Kedro的功能,具体步骤如下:
1. 自定义Markdown批量读取数据集
在你的Kedro项目的src/<project_name>/datasets/目录下新建Python文件(比如markdown_batch.py),编写自定义数据集类:
from pathlib import Path from typing import Dict from kedro.io import AbstractDataSet class MarkdownBatchDataSet(AbstractDataSet): def __init__(self, filepath: str): self._filepath = Path(filepath) def _load(self) -> Dict[str, str]: # 遍历目录下所有.md文件并读取内容 md_files = list(self._filepath.glob("*.md")) md_content = {} for file in md_files: with open(file, "r", encoding="utf-8") as f: md_content[file.name] = f.read() return md_content def _save(self, data: Dict[str, str]) -> None: # 若不需要保存功能,直接抛出未实现异常 raise NotImplementedError("MarkdownBatchDataSet不支持保存操作") def _describe(self) -> Dict[str, str]: return {"filepath": str(self._filepath)}
2. 在Catalog中配置自定义数据集
打开conf/base/catalog.yml,添加如下配置,将data/01_raw/folder_name/路径下的所有.md文件统一用raw_markdown_files名称管理:
raw_markdown_files: type: <project_name>.datasets.markdown_batch.MarkdownBatchDataSet filepath: data/01_raw/folder_name/
注意把<project_name>替换成你实际的Kedro项目名称。
3. 在Pipeline中调用数据集
你可以在Pipeline节点里直接使用这个数据集,示例代码如下:
from kedro.pipeline import Pipeline, node from .nodes import process_markdown_files def create_pipeline(**kwargs) -> Pipeline: return Pipeline( [ node( func=process_markdown_files, inputs="raw_markdown_files", outputs="processed_markdown", name="process_markdown_node", ) ] )
对应的节点处理函数示例:
def process_markdown_files(md_files: Dict[str, str]) -> Dict[str, dict]: processed = {} for filename, content in md_files.items(): # 这里写你的业务处理逻辑,比如提取标题、解析内容 lines = content.split("\n") title = next((line.strip("# ") for line in lines if line.startswith("#")), filename) processed[filename] = {"title": title, "content": content} return processed
这样就能通过raw_markdown_files这个Catalog名称,一次性获取指定目录下所有.md文件的内容了。
内容的提问来源于stack exchange,提问作者Lucas Jurani
相关产品推荐
相关产品推荐

