You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

API中本地Markdown转JSON遇语法错误,求解决方案

问题

我在API中尝试将本地Markdown文件转换为JSON时出现语法错误。起初以为读取文件存储为字符串变量时丢失了md2json所需的换行,但控制台打印变量时显示与原文件一致,仅在页面响应中显示无换行内容。后续测试发现,无换行的Markdown字符串可被md2json正常解析,带换行的反而不行。

使用的核心代码:

const md2json = require('md-2-json');
const fs = require('fs');

router.get('/page', (req, res) => {
    let content = fs.readFileSync('directory/file.md','utf8')
    console.log(content)
    let newContent = md2json.parse(content)
    console.log(newContent)
    res.send(newContent)
})

测试验证情况:

  • 可行代码(无换行的Markdown字符串):
let mdInput = `# Air ## The second largest heading ###### The smallest heading`

router.get('/page', (req, res) => {
    let newContent = md2json.parse(mdInput)
    console.log(newContent)
    res.send(newContent)
})
  • 不可行代码(带换行的Markdown字符串):
let mdInput = `
# Air

## The second largest heading
 
##### The smallest heading
`

router.get('/page', (req, res) => {
    let newContent = md2json.parse(mdInput)
    console.log(newContent)
    res.send(newContent)
})

请问如何解决该问题,或有无更优实现方式?


解决方法

1. 适配换行符并清理无效空白

md2json对换行符的兼容性较差,尤其是Windows环境下的\r\n换行,或者开头的空行可能触发解析bug。可以对读取的内容做预处理:

const md2json = require('md-2-json');
const fs = require('fs');

router.get('/page', (req, res) => {
    let content = fs.readFileSync('directory/file.md','utf8')
    // 统一转换为LF换行,去除开头空行
    content = content.replace(/\r\n/g, '\n').replace(/\r/g, '\n').trimStart();
    console.log(content)
    let newContent = md2json.parse(content)
    console.log(newContent)
    res.send(newContent)
})

2. 替换为更稳定的Markdown转JSON方案

md2json的维护状态一般,建议改用更成熟的库自定义转换逻辑,比如用markdown-it解析Markdown为Token,再生成需要的JSON结构:

const markdownIt = require('markdown-it');
const fs = require('fs');
const mdParser = markdownIt();

router.get('/page', (req, res) => {
    let content = fs.readFileSync('directory/file.md','utf8')
    // 解析Markdown得到Token数组
    const tokens = mdParser.parse(content, {});
    // 遍历Token生成自定义JSON结构
    const jsonResult = {};
    let currentHeading = '';

    tokens.forEach((token, index) => {
        // 匹配标题开始标签
        if (token.type === 'heading_open') {
            // 下一个Token是标题内容
            currentHeading = tokens[index + 1].content;
            jsonResult[currentHeading] = '';
        }
        // 匹配标题下的文本内容
        else if (token.type === 'inline' && currentHeading) {
            jsonResult[currentHeading] = token.content;
        }
    });

    console.log(jsonResult);
    res.send(jsonResult);
})

这种方式灵活性更高,能完全自定义输出的JSON结构,也不会受换行符或特殊格式的影响。

3. 排查文件本身的问题

确认本地Markdown文件的编码为UTF-8,没有隐藏的特殊字符(比如零宽空格、全角空格)。可以用VS Code等编辑器打开文件,查看右下角的编码标识,必要时重新保存为UTF-8编码。


内容的提问来源于stack exchange,提问作者Pontus Olsson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 14:20:35