PHP正则提取网页产品套餐及价格的技术求助
Hey there! Let's work through this price extraction problem you're stuck on. I totally get why regex might have failed before—dealing with inconsistent HTML where some package types (Single, 2-PACK, 4-PACK) might be missing can throw off rigid pattern matches. Let's build a solution that handles all cases and gives you the [type][price] array you need.
First: A Robust Regex Approach
Instead of trying to match all packages in one go (which breaks if any are missing), we'll target each package type individually. This way, even if a type isn't present on the page, it just stays as null in our result array instead of breaking the whole extraction.
Here's the PHP code:
<?php // Replace this with your actual HTML content (from cURL/file_get_contents etc.) $html = ' <div class="product-options"> <div class="option"> <h4>Single</h4> <span class="price-tag">$19.99</span> </div> <div class="option"> <h4>2-PACK</h4> <span class="price-tag">$34.99</span> </div> </div> '; // Initialize our result array with all possible types set to null (not found) $packagePrices = [ 'Single' => null, '2-PACK' => null, '4-PACK' => null ]; // Match Single package price if (preg_match('/(Single).*?(\$\d+\.\d{2})/is', $html, $singleMatches)) { $packagePrices[$singleMatches[1]] = $singleMatches[2]; } // Match 2-PACK package price if (preg_match('/(2-PACK).*?(\$\d+\.\d{2})/is', $html, $twoPackMatches)) { $packagePrices[$twoPackMatches[1]] = $twoPackMatches[2]; } // Match 4-PACK package price if (preg_match('/(4-PACK).*?(\$\d+\.\d{2})/is', $html, $fourPackMatches)) { $packagePrices[$fourPackMatches[1]] = $fourPackMatches[2]; } // Output the result print_r($packagePrices); ?>
Regex Breakdown:
(Single): Captures the package type as the first match group (we use this to map directly to our array key).*?: Non-greedy match for any characters (prevents it from skipping over other packages to find a price later in the HTML)(\$\d+\.\d{2}): Captures the price—matches a$, followed by one or more digits, a decimal point, and exactly two cents digits/ismodifiers:imakes the match case-insensitive (handles cases likesingleorSINGLE),sallows.to match newlines (critical since HTML elements are often split across lines)
If your HTML has specific classes or tags for package types/prices, you can make the regex more precise to avoid false matches. For example, if types are in <span class="package-type"> and prices in <span class="package-price">:
preg_match('/<span class="package-type">(Single)<\/span>.*?<span class="package-price">(\$\d+\.\d{2})<\/span>/is', $html, $matches)
A More Reliable Alternative: DOM Parsing
Regex works great for simple HTML, but if the page structure ever changes (even slightly), your regex might break. For more stability, use PHP's built-in DOMDocument and DOMXPath to parse the HTML properly:
<?php $html = ' <div class="product-options"> <div class="option"> <h4>4-PACK</h4> <span class="price-tag">$74.99</span> </div> </div> '; // Initialize result array $packagePrices = [ 'Single' => null, '2-PACK' => null, '4-PACK' => null ]; // Set up DOM parser (ignore minor HTML errors) $dom = new DOMDocument(); libxml_use_internal_errors(true); $dom->loadHTML($html); libxml_clear_errors(); $xpath = new DOMXPath($dom); // Check for each package type foreach ($packagePrices as $type => &$price) { // Look for the type text, then grab the following price element $priceNodes = $xpath->query("//*[contains(text(), '$type')]/following-sibling::*[contains(@class, 'price-tag')]"); if ($priceNodes->length > 0) { $price = trim($priceNodes->item(0)->nodeValue); } } print_r($packagePrices); ?>
This method navigates the HTML like a tree, so small changes (like extra whitespace or new elements) won't break it. It's the better choice for production code where HTML might evolve over time.
Either way, you'll end up with an array where each key is the package type, and the value is either the price string or null if that package isn't present on the page.
内容的提问来源于stack exchange,提问作者Ryan A

