Bash脚本中提取网页图片链接ID的正则表达式求助
Hey there! Let's work through extracting those image IDs from your fetched HTML in Bash. First, a quick syntax fix for your variable setup—Bash doesn't allow spaces around the equals sign when assigning variables, and quoting your URL avoids issues with special characters:
currentURL="https://www.example.com" currentPage=$(wget -q -O - "$currentURL")
情况1:提取<img>标签的id属性值
If you're targeting the id attribute directly on <img> elements (like <img id="hero-banner" src="...">), this Perl-compatible regex with grep will do the trick:
echo "$currentPage" | grep -oP '(?<=<img[^>]*id=")[^"]+'
Let's break this down:
(?<=<img[^>]*id="): A positive lookbehind that finds an<imgtag, skips any attributes beforeid=", and stops right at the opening quote of the ID value.[^"]+: Matches every character until the closing quote—this is your actual ID.
This handles those random &l... HTML entities you mentioned too, since it only cares about the content inside the quotes for the id attribute.
情况2:提取图片URL中的ID参数
If your "image link ID" refers to a query parameter in the image's src (like <img src="/photos?img_id=1234&size=large">), use this adjusted regex:
echo "$currentPage" | grep -oP '(?<=<img[^>]*src="[^?]*img_id=)[^&"]+'
Here's what's happening:
(?<=<img[^>]*src="[^?]*img_id=): Looks for thesrcattribute in an<img>tag, skips the base URL, and stops right afterimg_id=.[^&"]+: Grabs everything until the next&(other parameter) or closing quote—your target ID.
兼容单双引号的通用版本
If some tags use single quotes for attributes (like <img id='sidebar-photo' ...>), use this flexible regex that works with both quote types:
echo "$currentPage" | grep -oP '(?<=<img[^>]*id=([\'"]))(?(1)[^\1]+)'
A Quick Note
Regex works great for simple HTML, but for complex, nested pages, tools like pup (a command-line HTML parser) or xmllint are more reliable. But since you asked for regex, the above should cover your use case!
内容的提问来源于stack exchange,提问作者Mr. C

