Playwright爬取Kick.com遇Cloudflare检测,Cookie/UA无法绕过
问题:绕过Kick.com的Cloudflare机器人检测实现无限滚动爬取
我尝试使用Playwright爬取带有无限滚动功能的https://kick.com/browse/categories页面,执行了滚动JS代码并等待较长加载时间。关闭headless模式后,浏览器中页面可滚动几次,但滚动到底部加载更多内容时,请求返回403错误,提示未启用JavaScript或Cookie;开启headless模式后结果一致。
补充说明:我发现响应HTML的标题为“Just a moment...”,这是Cloudflare页面的特征,说明脚本被检测为机器人,正尝试寻找合适的Cookie/请求头以绕过检测。
请求详情:
Request: GET https://kick.com/api/v1/subcategories?limit=32&page=2 {'cluster': 'v1', 'sec-ch-ua-platform': '"Windows"', 'referer': 'https://kick.com/browse/categories', 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36', 'accept': 'application/json', 'sec-ch-ua': '"HeadlessChrome";v="129", "Not=A?Brand";v="8", "Chromium";v="129"', 'sec-ch-ua-mobile': '?0'}
我已尝试以下方法:
- 通过
page.set_extra_http_headers添加从Chrome开发者工具Network请求头中获取的Cookie - 将请求头中的
sec-ch-ua替换为Chrome中的对应值 - 使用playwright-stealth,但启用后目标页面加载为空白
请问还有哪些方法可以绕过该机器人检测?
我的代码
import asyncio from playwright.async_api import async_playwright import time # copied with cookie editor when opening the page with chrome cookies = [ { "domain": "kick.com", "expirationDate": 1726721839.170133, "name": "KP_UIDz-ssn", "path": "/", "value": "02drFYb8SXDKZm1tttHbDxrjgNuDdR4yBLOBCmEN9sCjbepNK6vZ1ESDUhPkXwwGDbuhype6dcvmnxsMesfoHNevNIZD4Htf8KWlaDjuN30u2N6SIIJciTgkEWG5nX8cWHWrf5qXLq8NU1SGriiT6yIfuTYvqEG3fvO1Kb" }, { "domain": ".kick.com", "name": "__cf_bm", "path": "/", "value": "CmeY2iVdWeXaYGUuIDyRMRxDgMQjYhZcl3we_Qy8pW8-1726689886-1.0.1.1-mLpIWwhnp3_zmFV7.bmPiIafI7q_BdbQduJiSxKUcOOCFPV.3r3Yb.p36MT3Sa1Ubq2TCTVOuilQEV3X6u0kjw" }, { "domain": ".kick.com", "name": "__stripe_mid", "path": "/", "value": "50b6ac3c-0cc6-44de-b64a-916706484819d40fd7" }, { "domain": ".kick.com", "name": "__stripe_sid", "path": "/", "value": "2d4b8647-0a68-46f9-8921-e8ce26409063280566" } ] PAGE_DOWN_JS = """ div = document.querySelector('#main-container'); div.scrollBy(0, window.innerHeight); """ async def handle_response(response): # Get the response body as text html_content = await response.body() print(f"Response: {html_content.decode('utf-8')}") async def main(): async with async_playwright() as p: # Launch the browser browser = await p.chromium.launch(headless=True, args=["--start-maximized"]) context = await browser.new_context(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36") await context.add_cookies(cookies) page = await context.new_page() # Navigate to the website await page.goto('https://kick.com/browse/categories') await page.wait_for_selector("#main-container") button = page.locator("div.flex.flex-row.items-center.justify-center.gap-2>button>>nth=0") await button.click() await page.screenshot(path="/Users/gxsong/kick/screenshot_before.png") page.on("console", lambda msg: print(f"Console [{msg.type}]: {msg.text}")) page.on("response", lambda response: asyncio.create_task(handle_response(response))) # Evaluate some JavaScript code on the page await page.evaluate(''' window.addEventListener('scroll', function() { console.log('Scrolled! Current scroll position:', window.scrollY); }); ''') for i in range(10): await page.evaluate(PAGE_DOWN_JS) time.sleep(10) await page.screenshot(path="/Users/gxsong/kick/screenshot.png") page_content = await page.content() # print(page_content) await browser.close() asyncio.run(main())
内容的提问来源于stack exchange,提问作者Ginni Song
相关产品推荐
相关产品推荐

