如何高效查找与目标URL匹配的正确oEmbed提供者?
我偶然了解到oEmbed规范,发现它有一个providers.json文件,里面是所有已知oEmbed提供者的信息,本质是一个大型数组,每个对象结构如下:
Vimeo示例
{ "provider_name": "Vimeo", "provider_url": "https://vimeo.com/", "endpoints": [ { "schemes": [ "https://vimeo.com/*", "https://vimeo.com/album/*/video/*", "https://vimeo.com/channels/*/*", "https://vimeo.com/groups/*/videos/*", "https://vimeo.com/ondemand/*/*", "https://player.vimeo.com/video/*" ], "url": "https://vimeo.com/api/oembed.{format}", "discovery": true } ] }
YouTube示例
{ "provider_name": "YouTube", "provider_url": "https://www.youtube.com/", "endpoints": [ { "schemes": [ "https://*.youtube.com/watch*", "https://*.youtube.com/v/*", "https://youtu.be/*", "https://*.youtube.com/playlist?list=*", "https://youtube.com/playlist?list=*", "https://*.youtube.com/shorts*" ], "url": "https://www.youtube.com/oembed", "discovery": true } ] }
我想在JavaScript项目里用这个文件,但不确定怎么高效实现:写一个接收URL的函数,找到匹配的提供者(如果有的话)。暴力解法是遍历每个提供者,把schemes里的每个条目转成正则表达式测试,直到找到匹配项,但这种方法效率低。有没有提速的方法?比如有没有比正则更高效的通配符匹配方式?
提速方案
1. 预构建域名索引,缩小匹配范围
oEmbed的scheme大多以域名区分,先把所有提供者按域名分组:
- 解析每个scheme的域名(比如
https://vimeo.com/*的域名是vimeo.com,https://*.youtube.com/watch*的域名是youtube.com) - 构建一个索引对象,键是域名,值是对应提供者的列表
这样处理输入URL时,先提取它的域名,直接从索引里拿到候选提供者,不用遍历全部。比如输入https://youtu.be/xxx,提取域名youtu.be,直接找对应YouTube的条目,跳过其他所有提供者。
2. 优化通配符匹配,避免全量正则转换
oEmbed的scheme里的通配符规则很简单,只有两种:*(匹配任意字符,包括路径分隔符)和*.(匹配任意子域名),可以自己写轻量的匹配函数,比正则表达式更快:
function matchScheme(url, scheme) { // 处理子域名通配符 if (scheme.startsWith('https://*.')) { const domainPart = scheme.slice(9); const urlDomain = new URL(url).hostname; if (!urlDomain.endsWith(domainPart)) return false; // 去掉域名部分,匹配剩余路径 const schemePath = scheme.slice(9 + domainPart.length); const urlPath = url.slice(url.indexOf(urlDomain) + urlDomain.length); return matchWildcardPath(urlPath, schemePath); } else { // 精确域名的情况,先匹配域名 const schemeUrl = new URL(scheme.replace('*', '')); const urlObj = new URL(url); if (schemeUrl.hostname !== urlObj.hostname) return false; // 匹配路径部分 return matchWildcardPath(urlObj.pathname + urlObj.search, schemeUrl.pathname + schemeUrl.search); } } function matchWildcardPath(path, pattern) { // 分割通配符前后片段,做前缀/后缀/中间匹配 const parts = pattern.split('*'); if (parts.length === 1) return path === pattern; // 前缀匹配 if (parts[0] && !path.startsWith(parts[0])) return false; // 后缀匹配 if (parts[parts.length - 1] && !path.endsWith(parts[parts.length - 1])) return false; // 中间片段依次匹配 let currentPos = parts[0].length; for (let i = 1; i < parts.length - 1; i++) { const part = parts[i]; const index = path.indexOf(part, currentPos); if (index === -1) return false; currentPos = index + part.length; } return true; }
这种字符串分割匹配的方式,避免了正则引擎的初始化和编译开销,在多数场景下比正则更快。
3. 预编译正则表达式(如果一定要用正则)
如果还是想用正则,不要每次匹配都重新编译,提前把所有scheme转成正则表达式缓存起来:
- 初始化时,遍历每个提供者的scheme,把
*转成.*,转义特殊字符(比如.转成\.),编译成RegExp对象并缓存 - 匹配时直接用预编译好的正则测试,省去重复编译的开销
4. 优先匹配高频提供者
如果你的场景里某些提供者(比如YouTube、Vimeo)出现频率很高,可以把它们放在索引的最前面,或者单独做一个快速匹配分支,先检查这些高频条目,命中后直接返回,不用遍历其他提供者。
总结
最有效的组合方式是:预构建域名索引 + 轻量通配符字符串匹配,既能大幅减少需要检查的提供者数量,又能避免正则的额外开销,在大多数场景下能把匹配速度提升数倍。
内容的提问来源于stack exchange,提问作者Svish

