You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP正则去除末尾字符串遇问题,寻求技术指导

Troubleshooting Regex for Mixed HTML/Text Content

Hey there! Let’s tackle your regex problem step by step. I get why you’re avoiding DOM parsing—when your content isn’t consistently <a> or <img> tags, regex is a reasonable approach (even if HTML regex has its quirks!).

First, let’s look at the example HTML you shared:

dscn8900

I’ll assume your goal is either extracting URLs from attributes like href, src, srcset, or modifying them (like updating domains). Here’s how to approach both scenarios:

1. Extracting URLs

For Tag Attributes

Use these regex patterns to capture values from common URL-bearing attributes:

  • Extract href values: href=['"]([^'"]+)['"]
    • This matches href=" or href=', then captures all characters until the closing quote.
  • Extract src values: src=['"]([^'"]+)['"]
    • Same logic as above, targeted at image sources.
  • Extract srcset values: srcset=['"]([^'"]+)['"]
    • Since srcset often has multiple URLs separated by commas, you’ll need to split the captured value afterward to get individual URLs.

For Plain Text URLs

If your content has raw URLs (not inside tags), use this pattern to catch them:
https?:\/\/[^\s<>"]+

  • Matches http:// or https://, then captures all characters until a whitespace, <, >, or quote.

Example Code (JavaScript)

Here’s how to put this into practice to extract all URLs from your content:

const content = '<a href="http://domain.co.uk.co.uk/wp-content/uploads/2016/06/DSCN8900.jpg"><img class="alignnone size-medium wp-image-4181" src="http://domain.co.uk/wp-content/uploads/2016/06/DSCN8900-300x225.jpg" alt="dscn8900" width="300" height="225" srcset="http://domain.co.uk/wp-content/uploads/2016/06/DSCN8900-1024x768.jpg 1024w, http://domain.co.uk/wp-content/uploads/2016/06/DSCN8900-300x225.jpg 300w"> Some plain text URL: http://domain.co.uk/another-page.html';

const extractedUrls = [];

// Extract from attributes
const attrRegex = /(href|src|srcset)=['"]([^'"]+)['"]/g;
let attrMatch;
while ((attrMatch = attrRegex.exec(content)) !== null) {
  const attrName = attrMatch[1];
  if (attrName === 'srcset') {
    // Split srcset into individual URLs
    const srcsetUrls = attrMatch[2].split(',').map(item => item.trim().split(' ')[0]);
    extractedUrls.push(...srcsetUrls);
  } else {
    extractedUrls.push(attrMatch[2]);
  }
}

// Extract plain text URLs
const plainRegex = /https?:\/\/[^\s<>"]+/g;
let plainMatch;
while ((plainMatch = plainRegex.exec(content)) !== null) {
  if (!extractedUrls.includes(plainMatch[0])) {
    extractedUrls.push(plainMatch[0]);
  }
}

console.log(extractedUrls);

2. Modifying URLs (e.g., Updating Domains)

If you need to replace parts of URLs (like changing domain.co.uk to new-domain.com), use a regex with capture groups:

const modifiedContent = content.replace(/(href|src|srcset)=['"](https?:\/\/)domain\.co\.uk([^'"]+)['"]/g, '$1="$2new-domain.com$3"');

This preserves the attribute name, protocol, and path while swapping out the domain.

Key Notes to Avoid Headaches

  • Quotation Marks: The patterns above support both single and double quotes. If your content uses unquoted attributes (rare but possible), adjust the regex to (href|src|srcset)\s*=\s*(['"]?)([^'"\s>]+)\1 to handle that case.
  • Edge Cases: Regex isn’t perfect for HTML—if your content has nested tags, escaped quotes, or malformed markup, you might need to tweak the patterns. For super messy content, a lightweight HTML parser might still be worth considering, but regex works great for structured, predictable content.

Let me know if you have a specific end goal (like replacing URLs, extracting specific types, etc.) and I can refine this further!

内容的提问来源于stack exchange,提问作者YaBCK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:32:24