如何绕过Python Selenium undetected_chromedriver窗口聚焦加载表格行限制
问题
我用undetected_chromedriver配合Selenium爬取页面 https://coinmarketcap.com/nft/collections/?page=1,用这段代码循环向下滚动加载内容:
driver.execute_script("window.scrollTo(0, {}-200)".format(pagescrollindex))
发现一个问题:浏览器窗口处于前台聚焦状态时,表格行能正常加载填充数据;但只要切到其他窗口覆盖它,表格行就变成空的。想问问有没有人碰到过这个情况,怎么绕开这个JS限制?
附上相关代码片段:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager import time import mysql.connector from random_user_agent.user_agent import UserAgent from random_user_agent.params import SoftwareName, OperatingSystem import undetected_chromedriver as uc index=0 while True: options = uc.ChromeOptions() options.headless = False driver = uc.Chrome(use_subprocess=True, options=options) time.sleep(4) try: print("get page:","https://coinmarketcap.com/nft/collections/?page={}".format(index)) url="https://coinmarketcap.com/nft/collections/?page={}".format(index) driver.get(url) if index==lastpage: index=1 trcount=0 retrycount=0 pagescrollindex=500 while trcount<100: driver.execute_script("window.scrollTo(0, {}-200)".format(pagescrollindex)) pagescrollindex=pagescrollindex+1000 time.sleep(3) htmlpage=driver.page_source curcount=htmlpage.count('<tr>') print("current count:",curcount) if trcount!=curcount: trcount=curcount retrycount=0 continue if retrycount==5: pagescrollindex=-5000 print("scroll back up") driver.execute_script("window.scrollTo(0, {}-200)".format(pagescrollindex)) pagescrollindex=200 retrycount=0 #break else: retrycount=retrycount+1 continue print()
解决方案
这种情况是网站通过页面可见性API做了懒加载限制:当窗口失焦时,暂停非必要的资源加载逻辑。以下是几种可行的绕过方法:
1. 强制页面认为自身始终可见
在页面加载后执行JS,重写可见性检测相关的API,让网站误以为页面一直处于前台:
# 页面打开后立即执行这段代码 driver.execute_script(""" Object.defineProperty(document, 'visibilityState', {get: () => 'visible'}); Object.defineProperty(document, 'hidden', {get: () => false}); document.addEventListener('visibilitychange', e => e.stopImmediatePropagation(), true); """)
2. 改用元素定位式滚动
放弃直接滚动窗口,定位到页面已加载的最后一行表格,用scrollIntoView触发滚动,更贴近用户真实操作,不受窗口聚焦状态影响:
# 替换原有的window.scrollTo逻辑 while trcount<100: try: # 定位最后一个<tr>元素,滚动到它的位置 last_tr = driver.find_elements(By.TAG_NAME, 'tr')[-1] driver.execute_script("arguments[0].scrollIntoView({block: 'end'});", last_tr) except IndexError: # 初始无数据时,直接滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") pagescrollindex += 1000 time.sleep(3) # 后续计数逻辑保持不变 htmlpage=driver.page_source curcount=htmlpage.count('<tr>') print("current count:",curcount) # ... 剩余代码略
3. 优化无头模式配置
如果可以接受无头模式,配置参数让它模拟真实有头浏览器的行为,从根源避免窗口覆盖的问题:
options = uc.ChromeOptions() options.headless = True # 添加必要参数模拟真实环境 options.add_argument("--window-size=1920,1080") options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("--enable-javascript") options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") driver = uc.Chrome(use_subprocess=True, options=options)
4. 保持窗口聚焦
通过模拟点击页面的方式,强制浏览器保持窗口激活状态:
from selenium.webdriver.common.action_chains import ActionChains # 在滚动循环中加入这段代码 while trcount<100: # 模拟点击页面空白处,维持窗口聚焦 ActionChains(driver).move_by_offset(20, 20).click().perform() # 原有滚动代码 driver.execute_script("window.scrollTo(0, {}-200)".format(pagescrollindex)) # ... 剩余代码略
内容的提问来源于stack exchange,提问作者Asaf David
相关产品推荐
相关产品推荐

