PHP正则去除末尾字符串遇问题,寻求技术指导
Hey there! Let’s tackle your regex problem step by step. I get why you’re avoiding DOM parsing—when your content isn’t consistently <a> or <img> tags, regex is a reasonable approach (even if HTML regex has its quirks!).
First, let’s look at the example HTML you shared:
I’ll assume your goal is either extracting URLs from attributes like href, src, srcset, or modifying them (like updating domains). Here’s how to approach both scenarios:
1. Extracting URLs
For Tag Attributes
Use these regex patterns to capture values from common URL-bearing attributes:
- Extract
hrefvalues:href=['"]([^'"]+)['"]- This matches
href="orhref=', then captures all characters until the closing quote.
- This matches
- Extract
srcvalues:src=['"]([^'"]+)['"]- Same logic as above, targeted at image sources.
- Extract
srcsetvalues:srcset=['"]([^'"]+)['"]- Since
srcsetoften has multiple URLs separated by commas, you’ll need to split the captured value afterward to get individual URLs.
- Since
For Plain Text URLs
If your content has raw URLs (not inside tags), use this pattern to catch them:https?:\/\/[^\s<>"]+
- Matches
http://orhttps://, then captures all characters until a whitespace,<,>, or quote.
Example Code (JavaScript)
Here’s how to put this into practice to extract all URLs from your content:
const content = '<a href="http://domain.co.uk.co.uk/wp-content/uploads/2016/06/DSCN8900.jpg"><img class="alignnone size-medium wp-image-4181" src="http://domain.co.uk/wp-content/uploads/2016/06/DSCN8900-300x225.jpg" alt="dscn8900" width="300" height="225" srcset="http://domain.co.uk/wp-content/uploads/2016/06/DSCN8900-1024x768.jpg 1024w, http://domain.co.uk/wp-content/uploads/2016/06/DSCN8900-300x225.jpg 300w"> Some plain text URL: http://domain.co.uk/another-page.html'; const extractedUrls = []; // Extract from attributes const attrRegex = /(href|src|srcset)=['"]([^'"]+)['"]/g; let attrMatch; while ((attrMatch = attrRegex.exec(content)) !== null) { const attrName = attrMatch[1]; if (attrName === 'srcset') { // Split srcset into individual URLs const srcsetUrls = attrMatch[2].split(',').map(item => item.trim().split(' ')[0]); extractedUrls.push(...srcsetUrls); } else { extractedUrls.push(attrMatch[2]); } } // Extract plain text URLs const plainRegex = /https?:\/\/[^\s<>"]+/g; let plainMatch; while ((plainMatch = plainRegex.exec(content)) !== null) { if (!extractedUrls.includes(plainMatch[0])) { extractedUrls.push(plainMatch[0]); } } console.log(extractedUrls);
2. Modifying URLs (e.g., Updating Domains)
If you need to replace parts of URLs (like changing domain.co.uk to new-domain.com), use a regex with capture groups:
const modifiedContent = content.replace(/(href|src|srcset)=['"](https?:\/\/)domain\.co\.uk([^'"]+)['"]/g, '$1="$2new-domain.com$3"');
This preserves the attribute name, protocol, and path while swapping out the domain.
Key Notes to Avoid Headaches
- Quotation Marks: The patterns above support both single and double quotes. If your content uses unquoted attributes (rare but possible), adjust the regex to
(href|src|srcset)\s*=\s*(['"]?)([^'"\s>]+)\1to handle that case. - Edge Cases: Regex isn’t perfect for HTML—if your content has nested tags, escaped quotes, or malformed markup, you might need to tweak the patterns. For super messy content, a lightweight HTML parser might still be worth considering, but regex works great for structured, predictable content.
Let me know if you have a specific end goal (like replacing URLs, extracting specific types, etc.) and I can refine this further!
内容的提问来源于stack exchange,提问作者YaBCK


