You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从外部菜谱页面提取指定的application/ld+json格式菜谱数据?

Extract Recipe Data from application/ld+json in PHP

Got it, let's tackle this problem. Instead of scraping specific HTML elements (which can break if the page structure changes), targeting the Schema.org Recipe structured data (the application/ld+json tags) is way more reliable—perfect for your use case since you noted the ld-json position might vary across pages.

Here's a revised PHP script that specifically filters out the Recipe data from all ld+json tags on the page:

<?php
$dom = new DOMDocument;
libxml_use_internal_errors(true);
// Load the target page
$dom->loadHTMLFile('https://www.thecookingcat.com/recipes/fluffy-pancake.php');
libxml_clear_errors();

$xpath = new DOMXPath($dom);
// Find all application/ld+json script tags
$scriptTags = $xpath->query("//script[@type='application/ld+json']");

$recipeData = null;

foreach ($scriptTags as $tag) {
    // Decode the JSON content from the script tag
    $jsonContent = json_decode($tag->nodeValue, true);
    
    // Skip if JSON decoding fails
    if (json_last_error() !== JSON_ERROR_NONE) {
        continue;
    }
    
    // Handle cases where the JSON is an array (some pages have multiple schema entries in one tag)
    if (is_array($jsonContent)) {
        foreach ($jsonContent as $entry) {
            if (isset($entry['@type']) && $entry['@type'] === 'Recipe') {
                $recipeData = $entry;
                break 2; // Exit both loops once we find the recipe
            }
        }
    } else {
        // Single schema entry in the script tag
        if (isset($jsonContent['@type']) && $jsonContent['@type'] === 'Recipe') {
            $recipeData = $jsonContent;
            break;
        }
    }
}

if ($recipeData) {
    // Output the recipe as formatted JSON (ready for your database)
    echo json_encode($recipeData, JSON_PRETTY_PRINT);
} else {
    echo "No Recipe structured data found on the page.";
}
?>

Key Details:

  • Flexible Targeting: This script doesn't care where the ld+json tags are placed—it finds all of them and filters by the @type field (looking for "Recipe").
  • Handles Edge Cases: Some pages might have multiple schema entries in a single ld+json tag (as an array), so the code checks both single objects and arrays.
  • Error Resilient: Skips any invalid JSON entries to avoid breaking the script.
  • Ready for Database: The final $recipeData variable contains all the structured recipe data (title, ingredients, instructions, cook time, etc.) in a native PHP array—you can directly insert this into your database or process it further as needed.

This approach is far more maintainable than scraping specific HTML elements, since Schema.org structured data is designed for programmatic access and is less likely to change when the page's visual layout updates.

内容的提问来源于stack exchange,提问作者Iwan Ross

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:17:45