You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用preg_match_all提取含YouTube/Vimeo的iframe src属性遇阻求方案

How to Capture YouTube/Vimeo iframe src Attributes with preg_match_all

Hey Devin, let's get this sorted out. Capturing specific iframe srcs can be tricky if your regex doesn't account for all the edge cases (like different quote types, domain variations, or case sensitivity). Here are two reliable approaches:

1. Targeted Regex Solution

This regex will match iframe tags where the src contains YouTube or Vimeo domains (including common variations like youtu.be or player.vimeo.com), and capture the full URL (with parameters).

Regex Pattern

/<iframe[^>]+src=["'](https?:\/\/(?:www\.)?(?:youtube\.com|youtu\.be|vimeo\.com|player\.vimeo\.com)\/[^"']+)["'][^>]*>/i

Breakdown of the Pattern:

  • /i: Case-insensitive modifier (matches Youtube, VIMEO, etc.)
  • <iframe[^>]+: Matches the start of an iframe tag, skipping any attributes before src
  • src=["']: Matches the src attribute opening quote (supports both single and double quotes)
  • (https?:\/\/(?:www\.)?(?:youtube\.com|youtu\.be|vimeo\.com|player\.vimeo\.com)\/[^"']+): The capture group for valid URLs:
    • https?:\/\/: Matches http:// or https://
    • (?:www\.)?: Optional www. prefix (non-capturing group)
    • (?:youtube\.com|youtu\.be|vimeo\.com|player\.vimeo\.com): Matches all common YouTube/Vimeo domain variations
    • \/[^"']+: Matches the rest of the URL (path + parameters) until the closing quote
  • ["'][^>]*>: Matches the closing quote and the rest of the iframe tag

PHP Code Example

$html = '
<iframe src="https://www.youtube.com/embed/dQw4w9WgXcQ?autoplay=1&controls=0"></iframe>
<iframe src=\'https://youtu.be/abc123\'></iframe>
<iframe src="https://player.vimeo.com/video/987654321"></iframe>
<iframe src="https://example.com/non-video"></iframe>
';

$pattern = '/<iframe[^>]+src=["\'](https?:\/\/(?:www\.)?(?:youtube\.com|youtu\.be|vimeo\.com|player\.vimeo\.com)\/[^"\']+)["\'][^>]*>/i';
preg_match_all($pattern, $html, $matches);

// The valid src URLs are stored in $matches[1]
print_r($matches[1]);

2. More Reliable DOM Parsing (Alternative to Regex)

Regex can fail with messy HTML (e.g., attributes in weird order, line breaks in tags). Using PHP's DOMDocument is a more robust approach for parsing HTML:

PHP Code Example

$html = '
<iframe src="https://www.youtube.com/embed/dQw4w9WgXcQ?autoplay=1"></iframe>
<iframe src="https://vimeo.com/123456789"></iframe>
<iframe src="https://example.com/video"></iframe>
';

$dom = new DOMDocument();
// Suppress HTML parsing errors (common for malformed HTML)
libxml_use_internal_errors(true);
$dom->loadHTML($html);
libxml_clear_errors();

$validSrcs = [];
$iframes = $dom->getElementsByTagName('iframe');

foreach ($iframes as $iframe) {
    $src = $iframe->getAttribute('src');
    // Check if the src contains a YouTube/Vimeo domain
    if (preg_match('/(youtube\.com|youtu\.be|vimeo\.com|player\.vimeo\.com)/i', $src)) {
        $validSrcs[] = $src;
    }
}

print_r($validSrcs);

Why DOM is Better:

  • Handles malformed HTML gracefully
  • Doesn't break if attributes are reordered or split across lines
  • Easier to maintain if you need to add more domain variations later

内容的提问来源于stack exchange,提问作者devin M. arnold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:03:47