收到200 OK状态码但网站仍跳转,原因是什么?
解决cloudscraper访问URL时无法获取前端JS跳转目标的问题
问题原因
浏览器中的跳转是前端JavaScript触发的页面内跳转,而非HTTP 3xx重定向,因此Network面板只会显示初始请求的200响应,看不到跳转的HTTP请求记录。cloudscraper仅获取初始页面的HTML内容,不会自动执行页面中的JS代码完成跳转。
解决方案
方案1:正则提取跳转URL或参数
直接解析初始页面的HTML内容,通过正则匹配跳转语句或生成目标URL所需的参数,手动拼接目标地址后请求数据。
代码示例:
import cloudscraper import re # 初始化cloudscraper scraper = cloudscraper.create_scraper() # 请求初始URL initial_response = scraper.get("https://monroe.county-taxes.com/public/search/property_tax?search_query=1573108&redirect=1573108") # 尝试匹配直接的JS跳转语句 jump_url_match = re.search(r'window\.location(?:\.href)?\s*=\s*["\']([^"\']+)["\']', initial_response.text) if jump_url_match: target_path = jump_url_match.group(1) target_url = f"https://monroe.county-taxes.com{target_path}" else: # 若没有直接跳转语句,匹配生成目标URL的核心参数 parcel_match = re.search(r'parcel\s*=\s*["\']([^"\']+)["\']', initial_response.text) qid_match = re.search(r'qid\s*=\s*["\']([^"\']+)["\']', initial_response.text) if not (parcel_match and qid_match): print("未找到跳转相关参数或URL") exit() # 拼接目标URL target_url = f"https://monroe.county-taxes.com/public/real_estate/parcels/1573108/bills?parcel={parcel_match.group(1)}&qid={qid_match.group(1)}" # 请求目标页面获取数据 target_response = scraper.get(target_url) print(target_response.text)
方案2:用execjs执行页面JS代码获取跳转URL
如果页面的跳转逻辑复杂(比如参数是通过JS动态计算生成的),可以使用execjs库直接执行页面中的JavaScript代码,获取最终的跳转地址。
代码示例:
import cloudscraper import execjs import re scraper = cloudscraper.create_scraper() initial_response = scraper.get("https://monroe.county-taxes.com/public/search/property_tax?search_query=1573108&redirect=1573108") # 提取页面中的核心JS代码 script_match = re.search(r'<script>([\s\S]*?)</script>', initial_response.text) if not script_match: print("未找到页面中的JS代码") exit() js_code = script_match.group(1) # 注入获取跳转地址的逻辑 js_code += "\nvar targetUrl = window.location.href;" # 执行JS获取目标URL ctx = execjs.compile(js_code) target_url = ctx.eval("targetUrl") # 补全完整URL(如果返回相对路径) if not target_url.startswith("http"): target_url = f"https://monroe.county-taxes.com{target_url}" # 请求目标页面 target_response = scraper.get(target_url) print(target_response.text)
注意事项
- 若网站的JS逻辑有混淆或加密,正则提取可能失效,此时
execjs是更可靠的选择,但需要确保环境中安装了Node.js(execjs依赖Node.js执行JS代码)。 - 两种方案均无需使用Selenium等无头浏览器,能保证请求速度。
内容的提问来源于stack exchange,提问作者Firefly
相关产品推荐
相关产品推荐

