You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mongoose菜谱搜索应用:跨集合匹配食材生成关键词求助

解决Mongoose菜谱食材匹配关键词的问题

嘿,我明白你遇到的问题了——你的聚合管道没生效的核心原因是字段内容不匹配:你的Recipe里的ingredients是带用量的复合字符串(比如"100g 西红柿"),但FoodLibrary的name是纯食材名称(比如"西红柿"),直接用$lookup做精确匹配肯定找不到对应结果。下面我给你一步步拆解解决方案:

第一步:从带用量的食材字符串中提取纯食材名

首先得把食材字符串里的用量、单位去掉,只留下食材名称。这里分两种常见场景处理:

场景1:食材字符串格式固定(比如「数量+单位+食材名」)

如果你的食材都是类似"200g 鸡胸肉"、"3个鸡蛋"这种格式,可以用MongoDB的$regexFind来精准提取食材部分:

{
  $addFields: {
    ingredientMatch: {
      $regexFind: {
        input: "$ingredients",
        // 正则匹配:开头是数字(支持小数)+ 常见单位 + 食材名
        regex: /^\d+(\.\d+)?\s*(?:g|kg|ml|l|tbsp|tsp|个|份|片)\s*(.+)$/,
        options: "i" // 忽略大小写
      }
    }
  }
},
{
  $addFields: {
    // 提取正则捕获到的食材名部分,如果没匹配到格式就用原字符串尝试匹配
    ingredientName: {
      $cond: {
        if: "$ingredientMatch",
        then: "$ingredientMatch.match.1",
        else: "$ingredients"
      }
    }
  }
}

场景2:食材格式不固定

如果有些食材字符串没有明确的用量(比如"少许盐"),可以简化处理,直接用$split分割后取后面的内容(这种方式对带空格的食材名比如"意大利面"也有效):

{
  $addFields: {
    ingredientName: {
      // 按空格分割后,取从第1个元素开始的所有内容并拼接
      $reduce: {
        input: { $slice: [{$split: ["$ingredients", " "]}, 1, { $size: {$split: ["$ingredients", " "]}}] },
        initialValue: "",
        in: { $concat: ["$$value", " ", "$$this"] }
      }
    }
  }
}

第二步:用模糊匹配关联FoodLibrary

提取到纯食材名后,不能再用精确匹配的$lookup,需要用$expr结合正则或者文本索引来做模糊匹配:

方案A:正则匹配(适合小数据量)

在$lookup里用$regexMatch实现模糊匹配:

{
  $lookup: {
    from: "foodlibraries",
    let: { ingName: "$ingredientName" }, // 传递提取到的食材名
    pipeline: [
      {
        $match: {
          $expr: {
            $regexMatch: {
              input: "$name",
              regex: "$$ingName",
              options: "i"
            }
          }
        }
      },
      // 只保留需要的字段(比如名称和分类),减少返回数据
      { $project: { name: 1, category: 1, _id: 0 } }
    ],
    as: "keywords"
  }
}

方案B:文本索引(适合大数据量,性能更好)

如果你的foodlibraries集合数据量很大,正则匹配会很慢,建议先给name字段创建文本索引:

// 在Mongoose模型里创建索引,或者直接在MongoDB shell里执行
FoodLibrary.createIndex({ name: "text" })

然后在$lookup的管道里用文本搜索:

{
  $lookup: {
    from: "foodlibraries",
    let: { ingName: "$ingredientName" },
    pipeline: [
      {
        $match: {
          $text: { $search: "$$ingName" }
        }
      },
      { $project: { name: 1, category: 1, _id: 0, score: { $meta: "textScore" } } },
      // 按匹配度排序,取最相关的结果
      { $sort: { score: -1 } }
    ],
    as: "keywords"
  }
}

第三步:重新分组回原Recipe文档

因为你之前用了$unwind展开食材数组,最后需要把数据重新分组,恢复成原来的Recipe结构,同时整理keywords数组:

{
  $group: {
    _id: "$_id",
    title: { $first: "$title" },
    ingredients: { $push: "$ingredients" },
    // 收集每个食材对应的keywords
    keywords: { $push: "$keywords" }
  }
},
// 扁平化keywords数组,去掉空匹配的结果
{
  $addFields: {
    keywords: {
      $filter: {
        input: { $reduce: { input: "$keywords", initialValue: [], in: { $concatArrays: ["$$value", "$$this"] } } },
        cond: { $ne: ["$$this", {}] }
      }
    }
  }
}

完整聚合管道示例

把上面的步骤整合起来,完整的代码如下:

Recipe.aggregate([
  // 展开食材数组,保留空数组的情况
  { "$unwind": { "path": "$ingredients", "preserveNullAndEmptyArrays": true } },
  // 提取纯食材名
  {
    $addFields: {
      ingredientMatch: {
        $regexFind: {
          input: "$ingredients",
          regex: /^\d+(\.\d+)?\s*(?:g|kg|ml|l|tbsp|tsp|个|份|片)\s*(.+)$/,
          options: "i"
        }
      }
    }
  },
  {
    $addFields: {
      ingredientName: {
        $cond: {
          if: "$ingredientMatch",
          then: "$ingredientMatch.match.1",
          else: "$ingredients"
        }
      }
    }
  },
  // 模糊匹配FoodLibrary
  {
    $lookup: {
      from: "foodlibraries",
      let: { ingName: "$ingredientName" },
      pipeline: [
        {
          $match: {
            $expr: {
              $regexMatch: {
                input: "$name",
                regex: "$$ingName",
                options: "i"
              }
            }
          }
        },
        { $project: { name: 1, category: 1, _id: 0 } }
      ],
      as: "keywords"
    }
  },
  // 重新分组回原Recipe
  {
    $group: {
      _id: "$_id",
      title: { $first: "$title" },
      ingredients: { $push: "$ingredients" },
      keywords: { $push: "$keywords" }
    }
  },
  // 整理keywords数组
  {
    $addFields: {
      keywords: {
        $filter: {
          input: { $reduce: { input: "$keywords", initialValue: [], in: { $concatArrays: ["$$value", "$$this"] } } },
          cond: { $ne: ["$$this", {}] }
        }
      }
    }
  }
])

额外注意事项

  1. 正则规则调整:根据你实际的食材字符串格式,调整正则里的单位列表(比如添加盎司、杯等)。
  2. 同义词处理:如果遇到类似番茄和西红柿这种同义词,可以在FoodLibrary里加一个synonyms数组字段,匹配时同时检查name和synonyms。
  3. 测试分步执行:调试时可以分步运行聚合管道,先看ingredientName的提取结果是否正确,再检查$lookup的匹配结果。

内容的提问来源于stack exchange,提问作者Deniz M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:22:53