使用Selenium+BeautifulSoup爬取Load More按钮页面仅获初始内容的问题
解决诈骗ICO列表爬取仅获取初始内容的问题
问题核心
点击"Load More"按钮后按钮消失,但爬取到的仍是初始列表内容,本质是点击后未等待新内容加载完成就立即抓取页面源码,且该按钮可能需要多次点击才能加载完所有诈骗ICO条目。
修正后的代码
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 初始化driver(此处假设你已完成driver初始化) ICO_n = [] Ended_L = [] category_L = [] # 循环点击Load More按钮,直到按钮不存在 while True: try: # 等待按钮可见 python_button = WebDriverWait(driver, 20).until( EC.visibility_of_element_located((By.XPATH, "//button[@data-type='SCAM']")) ) # 滚动到按钮位置(确保点击范围在可视区域) driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", python_button) # 执行点击操作 driver.execute_script("arguments[0].click();", python_button) # 等待新内容加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div.ico-item.ii-scm")) ) # 短暂等待页面渲染完成 time.sleep(1) except: # 按钮不存在时退出循环,说明所有内容已加载 break # 所有内容加载完成后再解析页面 soup = BeautifulSoup(driver.page_source, "html.parser") for a in soup.findAll('div', attrs={'class': "ico-item ii-scm"}): name = a.find("span", attrs={'class':"ii-name"}) ended = a.find("span", attrs={'class':"ii-det"}) category = a.find("span", attrs={'class':"ii-ico-cat"}) ICO_n.append(name.text.strip() if name else "--") Ended_L.append(ended.text.strip() if ended else "--") category_L.append(category.text.strip() if category else "--")
关键修改说明
- 循环点击逻辑:处理多页加载场景,直到按钮无法定位为止,确保所有诈骗ICO条目都被加载
- 显式等待替代固定休眠:用
WebDriverWait判断元素状态,比time.sleep更稳定,避免因网络延迟导致的内容未加载完成 - 文本处理优化:添加
.strip()清理文本多余空格,简化空值判断逻辑
内容的提问来源于stack exchange,提问作者Khouloud Safi
相关产品推荐
相关产品推荐

