如何比较URI编码字符串?Python/JS测试及解决方案
URI编码等价性问题的踩坑与解决
最近在一家成长型初创公司面试时碰到了个URI编码的棘手问题,W3.org明明说两个URI是等价的,但直接用编码工具处理后却没法直接比较,折腾了一番才找到解决方案,分享给大家:
问题核心
W3.org明确指出以下两个URI是等价的:
url4 = "http://b93.com:80/~smith/home.html"
url5 = "http://b93.com/%7Esmith/home.html"
但不管用Python还是JavaScript的原生编码方法处理,得到的结果都不一样,这就导致没法直接判断它们的等价性。
原生编码工具的测试结果
Python 测试(使用urllib.parse.quote)
直接对两个URL进行编码,结果差异很明显:
import urllib.parse url4 = "http://b93.com:80/~smith/home.html" url5 = "http://b93.com/%7Esmith/home.html" print(urllib.parse.quote(url4)) # 输出: 'http%3A//b93.com%3A80/~smith/home.html' print(urllib.parse.quote(url5)) # 输出: 'http%3A//b93.com/%257Esmith/home.html'
JavaScript 测试(使用encodeURIComponent)
同样的,用JS的原生编码方法处理后,结果也不一致:
const p1 = encodeURIComponent("http://b93.com:80/~smith/home.html"); const p2 = encodeURIComponent("http://b93.com/%7Esmith/home.html"); console.log(p1); // 输出: http%3A%2F%2Fb93.com%3A80%2F~smith%2Fhome.html console.log(p2); // 输出: http%3A%2F%2Fb93.com%2F%257Esmith%2Fhome.html
解决方案:URL标准化
当时我一直在想怎么正确比较这些编码后的字符串,后来发现Node.js的normalize-url库能完美解决这个问题——它会将等价的URL标准化为统一的格式,自动处理编码差异、默认端口等细节:
const normalizeUrl = require('normalize-url'); const n1 = normalizeUrl("http://b93.com:80/~smith/home.html"); const n2 = normalizeUrl("http://b93.com/%7Esmith/home.html"); console.log(n1); // 输出: http://b93.com/~smith/home.html console.log(n2); // 输出: http://b93.com/~smith/home.html
经过标准化处理后,两个原本编码不同但等价的URL就完全一致了,问题顺利解决。
内容的提问来源于stack exchange,提问作者Richard Rublev
相关产品推荐
相关产品推荐

