如何编写匹配.m3u/.m3u8结尾链接的正则表达式?
Got it, let's tackle this regex problem for extracting .m3u/.m3u8 links from your player initialization code or other text. Here's a solid solution tailored to common scenarios:
Core Regex Patterns
Basic Version (Matches any standalone or wrapped .m3u/.m3u8 link)
This works for most cases, including links embedded in code or plain text:
https?://[^\s]+\.(m3u8?)\b
Breakdown of each part:
https?://: Matches bothhttp://andhttps://links (the?makes thesoptional)[^\s]+: Captures all characters until a whitespace is hit—perfect since URLs don't normally include spaces\.(m3u8?): Targets the file extension:\.escapes the literal dot, andm3u8?matches either.m3uor.m3u8(the?makes the8optional)\b: Word boundary to ensure we don't accidentally match partial extensions (like.m3u8xyz)
Advanced Version (For links wrapped in quotes, common in player code)
If your target links are enclosed in single or double quotes (e.g., src: "https://example.com/stream.m3u8"), use this pattern to avoid false positives:
["'](https?://[^\s"']+\.(m3u8?))["']
The ["'] at the start and end matches quote characters, and the parentheses capture only the actual link inside them.
Example Usage (Python)
Here's how you'd implement this in code to extract links from player initialization text:
import re # Sample player initialization code player_code = ''' const videoPlayer = new VideoPlayer({ primarySource: "https://cdn.streaming.com/live/main.m3u8", fallbackSource: 'https://backup.streaming.com/live/backup.m3u', posterImage: "https://example.com/poster.jpg" }); ''' # Extract all quoted .m3u/.m3u8 links extracted_links = re.findall(r'["'](https?://[^\s"']+\.(m3u8?))["']', player_code) # Get just the full links (the first group in each match) full_links = [link[0] for link in extracted_links] print(full_links) # Output: ['https://cdn.streaming.com/live/main.m3u8', 'https://backup.streaming.com/live/backup.m3u']
Quick Adjustments for Edge Cases
- Relative paths: If your links don't start with
http/https(e.g.,/assets/stream.m3u), replacehttps?://with[^\s"']+to match any non-whitespace/non-quote characters before the extension. - Special characters in URLs: The pattern already handles characters like
?,&, and=since[^\s]includes all non-whitespace characters.
内容的提问来源于stack exchange,提问作者dmxyler
相关产品推荐
相关产品推荐

