如何用Regex仅捕获带前置美元符号的商品图片URL
Let's fix your regex to only grab image URLs tied to product prices (marked by a $). Here's a straightforward approach tailored to your content:
Step 1: Updated Regular Expression
Use this regex to target only images that follow a dollar sign, with no HTML tags breaking the connection between the price and image:
(?<=\$[^<]+)img src='([^']+)'
Step 2: Breakdown of the Regex
Let's break down why this works for your use case:
(?<=\$[^<]+): This is a positive lookbehind that verifies two key details right before theimgtag:- A literal dollar sign
\$(we escape it because$is a special regex character for end-of-string matches) - Any sequence of characters that aren't an opening HTML bracket
[^<]+— this ensures the$is part of the same product text block before the image, avoiding false matches from unrelated$symbols elsewhere
- A literal dollar sign
img src='([^']+)': Your original pattern to capture the image URL inside single quotes, retained to extract the actual URL value
Step 3: Test It on Your Sample Content
For your provided HTML snippet:
$400 + free shipping</title> <link>https://www.dealnews.com/...</link> <description><img src='http://c.dlnws.com/image/upload/f_auto,t_large,q_auto/content/vdiy8a75wg8v7bo92dhq'
This regex will successfully capture only the URL http://c.dlnws.com/image/upload/f_auto,t_large,q_auto/content/vdiy8a75wg8v7bo92dhq, since it's directly linked to the preceding $400 price. Non-product images (with no nearby $) will be ignored entirely.
Quick Edge Case Note
If your content ever has HTML tags between the $ and the img tag, you can adjust the lookbehind to allow for that with (?<=\$[\s\S]+?)img src='([^']+)' — just be aware this might catch unintended images if a $ appears much earlier in the content, so test against your full dataset first!
内容的提问来源于stack exchange,提问作者Corey Robinson

