如何用PHP file_get_contents获取指定网页track标签的src属性值
<track kind="captions"> src Attribute with PHP's file_get_contents Alright, let's walk through how to pull that captions track URL using PHP. First, we'll fetch the page content with file_get_contents, then parse the HTML to grab the specific attribute you need.
Method 1: Use DOMDocument (Recommended)
This is the most reliable approach because HTML can be messy, and DOM parsers handle variations like extra whitespace, attribute order changes, etc., way better than regex.
// Fetch the HTML from your target website $html = file_get_contents('https://your-target-site-url.com'); // Replace with your actual URL // Set up DOMDocument to handle potentially imperfect HTML $dom = new DOMDocument(); libxml_use_internal_errors(true); // Suppress parsing warnings for messy HTML $dom->loadHTML($html); libxml_clear_errors(); // Clear any stored warnings // Use XPath to find the <track> tag with kind="captions" $xpath = new DOMXPath($dom); $captionsTrack = $xpath->query('//track[@kind="captions"]'); // Check if we found the track element and get the src attribute if ($captionsTrack->length > 0) { $srcUrl = $captionsTrack->item(0)->getAttribute('src'); echo $srcUrl; // Will output: https://rolled.oped.info/sub/jAghd9t8AB4/HfQZ32SovcY.vtt } else { echo "Couldn't find a captions track on the page."; }
Method 2: Quick Regex (Fragile, Not Recommended for Production)
If you need a super quick solution and are confident the HTML structure won't change, you can use a regex. But be warned—this will break if the site modifies the HTML even slightly (like adding another attribute before src).
$html = file_get_contents('https://your-target-site-url.com'); $regexPattern = '/<track kind="captions" src="([^"]+)"/'; if (preg_match($regexPattern, $html, $matches)) { echo $matches[1]; // Outputs the captions URL } else { echo "No captions track found."; }
Pro Tip: Some sites block requests from
file_get_contentsby default. If you run into issues, add a user-agent header to mimic a real browser:$context = stream_context_create([ 'http' => [ 'header' => 'User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' ] ]); $html = file_get_contents('https://your-target-site-url.com', false, $context);
内容的提问来源于stack exchange,提问作者Jacob K.

