You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Flutter中实现适配多网站的通用食谱提取功能?

网页食谱提取通用化方案及可行性分析

需求可行性

完全可行,已有成熟技术方案和开源项目验证了这类功能的落地性,核心是通过标准化识别逻辑覆盖不同网站的结构差异。

通用化实现步骤

1. 优先解析Schema.org结构化数据

多数正规食谱网站会嵌入Schema.org的Recipe类型结构化数据,这是最可靠的提取方式,无需依赖网站DOM结构:

  • 定位HTML中的<script type="application/ld+json">标签,直接读取标准化字段:name对应标题,recipeIngredient对应食材列表,recipeInstructions对应制作步骤
  • 示例代码片段:
import 'dart:convert';

void extractFromSchema(String htmlContent) {
  dom.Document document = parser.parse(htmlContent);
  var scriptTags = document.querySelectorAll('script[type="application/ld+json"]');
  
  for (var script in scriptTags) {
    try {
      var jsonData = jsonDecode(script.text);
      // 处理部分网站返回数组的情况
      if (jsonData is List) jsonData = jsonData.first;
      
      if (jsonData['@type'] == 'Recipe') {
        setState(() {
          _title = jsonData['name'] ?? "未找到标题";
          _ingredients = List<String>.from(jsonData['recipeIngredient'] ?? []);
          _directions = (jsonData['recipeInstructions'] as List?)
                  ?.map((step) => step['text'] ?? step.toString().trim())
                  .where((text) => text.isNotEmpty)
                  .toList() ?? [];
        });
        return; // 找到有效数据后直接终止解析
      }
    } catch (e) {
      continue; // 解析失败则跳过当前脚本标签
    }
  }
}

2. 维护多网站规则映射表

针对没有Schema数据的网站,建立域名与CSS选择器的映射表,覆盖主流食谱站点:

  • 创建规则表示例:
final Map<String, Map<String, String>> siteRules = {
  'allrecipes.com': {
    'title': 'h1',
    'ingredients': '.mntl-structured-ingredients__list-item',
    'directions': '#recipe__steps-content_1-0 .mntl-sc-block-html'
  },
  'foodnetwork.com': {
    'title': '.o-AssetTitle__a-HeadlineText',
    'ingredients': '.o-Ingredients__a-Ingredient',
    'directions': '.o-Method__m-Step'
  },
  // 可继续添加其他网站规则
};
  • 提取时先解析URL的域名,匹配对应规则进行内容抓取:
String getDomainFromUrl(String url) {
  Uri uri = Uri.parse(url);
  return uri.host.replaceAll('www.', '');
}

void extractWithSiteRules(String htmlContent, String domain) {
  if (!siteRules.containsKey(domain)) return;
  
  dom.Document document = parser.parse(htmlContent);
  var rules = siteRules[domain]!;
  
  setState(() {
    _title = document.querySelector(rules['title'])?.text.trim() ?? "未找到标题";
    
    _ingredients = document.querySelectorAll(rules['ingredients']!)
        .map((e) => e.text.trim())
        .where((text) => text.isNotEmpty)
        .toList();
        
    _directions = document.querySelectorAll(rules['directions']!)
        .map((e) => e.text.trim())
        .where((text) => text.isNotEmpty)
        .toList();
  });
}

3. 语义关键词 fallback 提取

如果前两种方式都失效,通过关键词定位目标内容:

  • 遍历页面中的标题、列表元素,搜索"食材"、"Ingredients"、"步骤"、"Directions"等关键词
  • 提取关键词相邻的列表或文本块,作为食材和步骤内容,适合小众或非结构化的食谱页面

4. 用户体验优化

  • 加载状态:请求和提取过程中显示加载动画,避免用户误解无响应
  • 错误提示:提取失败时给出明确提示,比如"无法识别该网页的食谱内容,请尝试其他链接"
  • 手动编辑:允许用户修改自动提取的标题、食材和步骤,弥补自动识别的不足

注意事项

  • 反爬规避:添加User-Agent请求头模拟浏览器访问,避免被网站拦截
  • 动态内容处理:若网站用JavaScript渲染内容,改用webview_flutter加载页面后提取渲染完成的HTML
  • 规则更新:定期维护网站规则表,适配站点结构的变更

内容的提问来源于stack exchange,提问作者Leena Marie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 05:39:54