如何用TypeScript正则表达式从URL中提取指定URN字符串?
问题
给定URL字符串:
https://test.io/content/storage/id/urn:aaid:sc:US:8eda16d4-baba-4c90-84ca-0f4c215358a1;revision=0?component_id=e62a5567-066d-452a-b147-19d909396132
需要从中提取以下字符串:
urn:aaid:sc:US:8eda16d4-baba-4c90-84ca-0f4c215358a1
该目标字符串始终以urn开头,以字母或数字结尾。尝试了以下TypeScript代码,但返回null:
function extractAssetIdFromUrl(url: string) { // Regular expression to match the desired pattern const regex = /urn[\w-]+/; // Use the regex to find the match in the URL const match = url.match(regex); // Check if there is a match and return it, otherwise return null return match ? match[0] : null; }
解决方案
原正则无法匹配的核心问题是:\w只包含字母、数字和下划线,不包含目标字符串里的冒号:,所以匹配到urn后遇到第一个:就会停止,无法覆盖完整的目标内容。
针对需求,调整正则规则即可解决,以下是两种可行的实现方式:
方式一:基于URL结构的精准匹配
利用URL中目标字符串后紧跟;或?的特点,用正向预查限制匹配范围:
function extractAssetIdFromUrl(url: string) { // 匹配urn开头,包含:、字母、数字、连字符的内容,直到遇到;或?为止 const regex = /urn:[a-zA-Z0-9:-]+(?=[;?])/; const match = url.match(regex); return match ? match[0] : null; }
正则说明:
urn::明确锁定起始标识,避免误匹配其他含urn的子串[a-zA-Z0-9:-]+:允许匹配目标字符串里的所有合法字符(字母、数字、冒号、连字符)(?=[;?]):正向预查,确保匹配内容的结尾紧跟URL里的分隔符,不会额外捕获后面的参数内容
方式二:基于结尾规则的通用匹配
如果目标字符串的结束位置不确定,直接确保结尾是字母或数字:
function extractAssetIdFromUrl(url: string) { // 匹配urn开头的合法字符,最终以字母或数字结尾 const regex = /urn:[a-zA-Z0-9:-]*[a-zA-Z0-9]/; const match = url.match(regex); return match ? match[0] : null; }
正则说明:
[a-zA-Z0-9:-]*:匹配任意数量的合法字符[a-zA-Z0-9]:强制结尾为字母或数字,符合需求定义
内容的提问来源于stack exchange,提问作者Rue Vitale
相关产品推荐
相关产品推荐

