如何用preg_match从file_get_contents返回的多content属性中提取媒体内容
使用preg_match提取meta标签中content属性的媒体内容
嘿,我来帮你搞定这个问题!你已经用file_get_contents拿到了目标页面的meta标签内容,现在要提取content属性里的媒体链接,用preg_match完全可以做到,分两种场景给你说明:
1. 精准提取视频媒体链接(针对og:video:secure_url)
如果你只需要那个MP4链接,直接针对og:video:secure_url属性写正则会更精准,不会拿到其他无关的content值。代码示例如下:
// 假设这是你获取到的HTML内容 $html = '<meta property="og:video:secure_url" content="//abc-def- ghi.net/reg02/2016/05/31/05/92f24c57-2cb3-42b0-84cb-edad24d7b68f.mp4? secure=1"><meta property="og:video:type" content="video/mp4"><meta property="og:video:width" content="540"><meta property="og:video:height" content="960"><meta property="fb:app_id" content="682659535154406">'; // 匹配目标meta标签的content值 if (preg_match('/meta\s+property="og:video:secure_url"\s+content="([^"]+)"/i', $html, $matches)) { $mediaUrl = $matches[1]; echo "提取到的媒体链接:" . $mediaUrl; }
正则细节说明:
\s+:匹配任意数量的空格,避免标签里属性之间空格不一致导致匹配失败([^"]+):捕获双引号之间的所有内容(除了双引号本身),这就是我们要的媒体链接/i:忽略大小写,兼容HTML标签可能出现的大小写差异(比如<META>或者Property)
2. 批量提取所有meta标签的content属性值
如果需要拿到所有content里的内容(比如视频类型、尺寸、FB应用ID),就用preg_match_all来批量提取:
preg_match_all('/content="([^"]+)"/i', $html, $allMatches); // $allMatches[1]就是所有content值组成的数组 print_r($allMatches[1]);
运行后你会得到一个数组,包含视频链接、video/mp4、540、960和FB应用ID这些内容。
额外兼容提示
如果目标页面的meta标签属性可能用单引号(比如content='xxx'),可以把正则调整得更灵活,同时兼容单双引号:
preg_match('/meta\s+property="og:video:secure_url"\s+content=([\'"])([^\1]+)\1/i', $html, $matches); $mediaUrl = $matches[2];
这个写法会自动识别属性使用的是单引号还是双引号,适配更多页面场景。
内容的提问来源于stack exchange,提问作者dudeiamstoned
相关产品推荐
相关产品推荐

