正则表达式在Regex101可用但PHP无匹配,请求技术协助
Alright, let's troubleshoot why your regex works on Regex101 but fails in PHP—this is a super common gotcha, so I’ll walk you through the most likely fixes:
Your regex uses < and > (HTML entities for < and >), but if you’re matching raw HTML content (not escaped entity strings) in PHP, these entities need to be replaced with actual angle brackets. Regex101 might have let you test with escaped entities, but PHP is looking for the real tags.
For example, change <div class="panel-heading"> to <div class="panel-heading"> in your pattern.
s (DOTALL) Modifier PHP’s regex engine defaults to having . not match newline characters. Most web HTML has line breaks between elements, so your (.+?) capture groups stop short at the first newline. Adding the s modifier at the end of your pattern makes . match all characters, including newlines—this matches the "dot matches newline" setting you probably used in Regex101.
Web HTML often has inconsistent whitespace (spaces, tabs, newlines) between tags, but your regex uses fixed spaces. Replace fixed spaces with \s* (matches any number of whitespace characters) to account for this.
Example Corrected Regex
$pattern = '/<div class="panel-heading"><a href="(.+?)"><h5>(.+?)<\/h5><\/a><\/div>\s*<div class="panel-body">\s*<p><b> Author: <\/b> (.+?)<\/p>\s*<p><b> Awarding University: <\/b> (.+?)<\/p>\s*<p><b> Level : <\/b> (.+?)<\/p>\s*<p><b> Year: <\/b> (.+?)<\/p>/s';
If you’re trying to match multiple panels on the page, preg_match will only return the first match. Use preg_match_all instead to capture all instances, and use PREG_SET_ORDER to organize matches into easy-to-use arrays:
Sample PHP Code
$html = file_get_contents('your-target-page.html'); // Or your raw HTML string if (preg_match_all($pattern, $html, $matches, PREG_SET_ORDER)) { foreach ($matches as $item) { echo "URL: " . $item[1] . "\n"; echo "Title: " . $item[2] . "\n"; echo "Author: " . $item[3] . "\n"; echo "University: " . $item[4] . "\n"; echo "Level: " . $item[5] . "\n"; echo "Year: " . $item[6] . "\n\n"; } } else { echo "No matches found."; }
Regex is fragile for HTML—small changes like extra attributes or tag nesting will break your pattern. PHP’s built-in DOMDocument and XPath are way more robust for web scraping:
$dom = new DOMDocument(); libxml_use_internal_errors(true); // Ignore minor HTML parsing errors $dom->loadHTML($html); libxml_clear_errors(); $xpath = new DOMXPath($dom); $panels = $xpath->query('//div[contains(@class, "panel-heading")]/parent::div'); foreach ($panels as $panel) { // Extract data with XPath queries $link = $xpath->query('.//div[@class="panel-heading"]/a', $panel)->item(0)->getAttribute('href'); $title = $xpath->query('.//div[@class="panel-heading"]/a/h5', $panel)->item(0)->nodeValue; $author = trim(str_replace('Author:', '', $xpath->query('.//div[@class="panel-body"]/p[b[text()=" Author: "]]', $panel)->item(0)->nodeValue)); echo "URL: $link\nTitle: $title\nAuthor: $author\n\n"; }
This method won’t break if the HTML adds extra spaces or rearranges attributes!
内容的提问来源于stack exchange,提问作者Ruben Alves

