正则表达式新手求助:使用Matcher提取网页特定URL失败
Hey there! Let's get that URL extraction working for you. The issue with your current pattern \"https:.*?\" is that it's way too broad—it might grab random HTTPS URLs you don't care about, and could miss your target URLs if there's unexpected whitespace or specific structure they follow.
Breaking Down Your Target URL
From your example, your desired URLs have a strict fixed structure with two variable parts (the bold sections):
" https://scontent-mad1-1.cdntest.com/vp/ 8e2f9af8362ea131062778fccbfda09e/5AEB6394/t51.2885-19/s320x320/ 12918047_263707850629206_1737689122_a.jpg"
Let's map the fixed vs variable parts:
- Fixed start:
" https://scontent-mad1-1.cdntest.com/vp/ - Variable part 1: Hexadecimal characters (0-9, a-f) — that long hash string
- Fixed middle:
/5AEB6394/t51.2885-19/s320x320/ - Variable part 2: Numbers and underscores — the ID-like string
- Fixed end:
_a.jpg"
The Correct Regex Pattern
We can build a precise regex to match only this structure, with capture groups for the variable parts:
" https://scontent-mad1-1.cdntest.com/vp/([0-9a-f]+)/5AEB6394/t51.2885-19/s320x320/([\\d_]+)_a.jpg"
Regex Breakdown:
" https://scontent-mad1-1.cdntest.com/vp/: Exact match for the quoted start, space, and fixed domain/path([0-9a-f]+): First capture group — matches one or more hex characters (for the hash)/5AEB6394/t51.2885-19/s320x320/: Exact match for the fixed middle path segment([\\d_]+): Second capture group — matches one or more digits/underscores (for the ID)_a.jpg": Exact match for the fixed quoted end
Java Code Example
Here's how to use this with Pattern and Matcher to extract your URLs:
import java.util.regex.Matcher; import java.util.regex.Pattern; public class UrlExtractor { public static void main(String[] args) { // Replace this with your actual HTML source String html = "Your full webpage source code here"; // The regex (note escaped double quotes and backslashes for Java strings) String regex = "\" https://scontent-mad1-1.cdntest.com/vp/([0-9a-f]+)/5AEB6394/t51.2885-19/s320x320/([\\d_]+)_a.jpg\""; Pattern pattern = Pattern.compile(regex); Matcher matcher = pattern.matcher(html); // Loop through all matches while (matcher.find()) { // Get the full matched URL (including quotes) String fullUrl = matcher.group(0); System.out.println("Matched URL: " + fullUrl); // Get the first variable part (the hash) String hashSegment = matcher.group(1); System.out.println("Hash segment: " + hashSegment); // Get the second variable part (the ID string) String idSegment = matcher.group(2); System.out.println("ID segment: " + idSegment); } } }
Quick Adjustments If Needed
- If the space after the opening quote isn't always there, replace the space with
\\s?(matches zero or one whitespace character):"\"\\s?https://scontent-mad1-1.cdntest.com/vp/([0-9a-f]+)/5AEB6394/t51.2885-19/s320x320/([\\d_]+)_a.jpg\"" - If the hash could include uppercase letters, update the first capture group to
[0-9a-fA-F]+
内容的提问来源于stack exchange,提问作者YASSBRA Top10

