求助:使用Selenium Python抓取带加载更多按钮的博客标题仅获第一页内容
问题根源
你代码里的核心问题是一直用初始请求的静态页面内容解析,完全没用到Selenium加载后的动态页面:
- 一开始用
requests.get(urls)获取的r.content是页面首次加载的静态源码,循环里反复解析这个内容,自然只会抓到第一页标题。 - 点击「Show more」后页面已经动态加载了新内容,但你没有重新获取当前浏览器的页面源码,还是用最开始的旧内容解析。
修正后的代码
from bs4 import BeautifulSoup import pandas as pd import time # Selenium Routine from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.keys import Keys # Removes SSL Issues With Chrome options = webdriver.ChromeOptions() options.add_argument('--ignore-certificate-errors') options.add_argument('--ignore-ssl-errors') options.add_argument('--ignore-certificate-errors-spki-list') options.add_argument('log-level=3') options.add_argument('--disable-notifications') #options.add_argument('--headless') # Comment to view browser actions # 初始化浏览器并打开目标页面 urls = "https://jooble.org/blog/" driver = webdriver.Chrome(executable_path="C:\webdrivers\chromedriver.exe", options=options) driver.get(urls) productlist = [] for i in range(1, 3): # 关键:每次循环都获取当前浏览器的页面源码,而不是初始静态内容 soup = BeautifulSoup(driver.page_source, features='lxml') items = soup.find_all('div', class_='post') print(f'LOOP: start [{len(items)}]') for single_item in items: title = single_item.find('div', class_='front__news-title').text.strip() print('Title:', title) productlist.append({'Title': title}) print() # 避免重复处理最后一页(如果循环到最后一次就不用点击了) if i < 2: time.sleep(2) # 可以缩短等待时间,不需要5秒 WebDriverWait(driver, 40).until( EC.element_to_be_clickable((By.XPATH, "//button[normalize-space()='Show more']")) ).send_keys(Keys.ENTER) driver.close() # Save Results df = pd.DataFrame(productlist) df.to_csv('Results.csv', index=False)
关键改动说明
- 删掉了无用的
requests.get(urls)和r变量,直接用driver.page_source获取当前浏览器的动态页面源码。 - 每次循环都重新生成
BeautifulSoup对象,确保解析的是最新加载的页面内容。 - 添加了判断逻辑,最后一次循环不再点击「Show more」,避免不必要的等待。
- 缩短了等待时间(从5秒改为2秒),提升效率,同时保留显式等待确保按钮可点击。
内容的提问来源于stack exchange,提问作者MarkWP
相关产品推荐
相关产品推荐

