使用Python Requests库爬取网站时如何绕过反广告拦截页面?
解决爬虫被拦截的方案
针对你遇到的网站返回广告拦截提示的问题,本质是网站的反爬机制识别出了你的请求并非来自真实浏览器,以下是几个可行的解决思路:
使用带JS渲染的爬虫工具
纯requests请求无法处理网站的JS检测逻辑,改用Selenium或Playwright这类能模拟完整浏览器行为的工具,它们可自动执行页面JS,绕过这类检测。示例代码(Selenium):from selenium import webdriver from selenium.webdriver.chrome.options import Options import time def get_zip_page(): chrome_options = Options() chrome_options.add_argument('--user-agent=Mozilla/5.0 (Linux; Android 6.0; Nexus 5 Build/MRA58N) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.107 Mobile Safari/537.36') # 禁用扩展,避免广告拦截插件干扰页面加载 chrome_options.add_argument('--disable-extensions') driver = webdriver.Chrome(options=chrome_options) driver.get("https://www.unitedstateszipcodes.org") # 等待页面完全加载 time.sleep(3) page_source = driver.page_source driver.quit() return page_source data = get_zip_page()补充完整请求头并携带有效Cookie
网站可能通过检查请求头完整性和Cookie验证请求合法性。你可以从真实浏览器访问该网站时,复制完整的请求头(包括Accept、Accept-Language、Referer、Cookie等字段)添加到requests请求中:import requests def get_data(link): hdr = { 'user-agent': 'Mozilla/5.0 (Linux; Android 6.0; Nexus 5 Build/MRA58N) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.107 Mobile Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.google.com/', # 从浏览器复制当前有效的Cookie内容 'Cookie': '此处替换为你的Cookie值' } req = requests.get(link, headers=hdr) return req.text data = get_data("https://www.unitedstateszipcodes.org")注意:Cookie存在过期时间,需要定期更新。
模拟人类行为节奏
在请求之间添加随机间隔(比如time.sleep(2-5)),避免短时间内频繁发起请求,降低被识别为爬虫的概率。考虑官方API替代方案
若追求数据准确性,USPS的官方地址验证API是更可靠的选择,它能直接返回标准邮政编码,相比爬取第三方网站更稳定,不过需要先申请API密钥。
内容的提问来源于stack exchange,提问作者snowball
相关产品推荐
相关产品推荐

