PHP中使用preg_match提取HTML文档lang属性值的正则匹配优化需求
Got it, let's fix those regex issues you're facing! Your current pattern has two main problems: it expects lang to be the first attribute in the <html> tag, and it doesn't properly limit the capture to just the language code when other attributes follow. Here are two solid solutions:
Solution 1: Improved Regular Expression
This regex will handle lang attributes anywhere in the <html> tag, and only capture valid language code formats (letters and hyphens, like en, de-DE, etc.):
$sitelang = preg_match('!<html\b[^>]*lang="([a-z-]+)"[^>]*>!i', $result, $matches) ? $matches[1] : 'Site Language not detected';
Breakdown of the regex:
!<html\b: Ensures we match the full<html>tag (the\bword boundary prevents matches for tags like<html5>)[^>]*: Matches any characters except>(covers all other attributes or whitespace before/after thelangattribute)lang="([a-z-]+)": Captures only valid language code characters (letters and hyphens) inside the escaped quotes. Using[a-z-]+instead of.*?avoids accidentally grabbing extra content from subsequent attributes.[^>]*>: Matches any remaining content in the<html>tag until the closing>!i: Case-insensitive modifier, so it works with<HTML>orLANG=
Test cases it handles:
<html lang="de">→ Returnsde<html lang="en" class="desktop-view not-mobile-device text-size-normal anon">→ Returnsen<html data-n-head-ssr lang="en">→ Returnsen<HTML LANG="en-GB" data-test="123">→ Returnsen-GB
Solution 2: Use DOMDocument (More Reliable for HTML)
Regex can struggle with edge cases like malformed HTML, line breaks in tags, or unusual attribute formatting. For a more robust approach, use PHP's built-in DOM parser:
// Initialize DOMDocument and suppress parsing errors (common for real-world HTML) $dom = new DOMDocument(); libxml_use_internal_errors(true); $dom->loadHTML($result); libxml_clear_errors(); // Get the <html> tag and extract the lang attribute $htmlTag = $dom->getElementsByTagName('html')->item(0); $sitelang = $htmlTag ? trim($htmlTag->getAttribute('lang')) : 'Site Language not detected';
This method doesn't care about attribute order or tag formatting—it directly accesses the lang attribute of the <html> element, which is the proper way to parse HTML content.
内容的提问来源于stack exchange,提问作者mem

