如何检测链接是否跳转到域外?排查Python跳转检测代码失效问题
问题排查与解决方案
你的代码无法正确检测跨域名跳转,核心问题如下:
- 逻辑错误:设置
allow_redirects=True时,requests会自动跟随所有重定向,最终返回的response.status_code是目标页面的200状态码,导致你的判断逻辑直接忽略所有跳转行为,误判为"Not"。 - 缺少域名对比核心逻辑:代码没有提取原始URL与最终跳转URL的域名进行对比,这是判断跨域跳转的关键。
- HEAD请求兼容性差:部分服务器会拒绝HEAD请求,导致无法获取正确的跳转信息。
- Location字段不可靠:多次跳转时,
response.headers['Location']仅记录最后一次跳转地址,且最终状态为200时该字段大概率不存在。
修正后的代码
import requests from urllib.parse import urlparse def check_external_redirect(original_url): try: # 用GET请求替代HEAD,提升兼容性,设置超时避免挂起 response = requests.get(original_url, allow_redirects=True, timeout=10) # 解析原始URL与最终跳转URL的域名 original_domain = urlparse(original_url).netloc final_domain = urlparse(response.url).netloc # 标准化域名(统一小写、去除www前缀,避免无意义差异) def normalize_domain(domain): return domain.lower().replace('www.', '') normalized_original = normalize_domain(original_domain) normalized_final = normalize_domain(final_domain) # 判断是否跨域名跳转 if normalized_original != normalized_final: return original_url, f"YES to external domain: {response.url}" else: # 区分无跳转和同域名跳转 if len(response.history) > 0: return original_url, f"Redirected but within same domain: {response.url}" else: return original_url, "No redirect" except requests.exceptions.RequestException as e: return original_url, f"Error: {str(e)}" # 测试目标URL url = "https://www.gazetadopovo.com.br/conteudo-publicitario/some-control-summit/daos-conheca-organizacao-baseada-em-redes-blockchain/" print(check_external_redirect(url))
关键改进说明
- 域名解析与对比:通过
urlparse提取域名,标准化后对比,确保准确识别跨域跳转。 - 跳转行为追踪:利用
response.history判断是否存在跳转,即使是同域名跳转也能明确识别。 - 兼容性优化:改用GET请求并设置超时,避免服务器拒绝请求或请求挂起。
内容的提问来源于stack exchange,提问作者Rafael
相关产品推荐
相关产品推荐

