You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Selenium Python抓取带加载更多按钮的博客标题仅获第一页内容

问题根源

你代码里的核心问题是一直用初始请求的静态页面内容解析,完全没用到Selenium加载后的动态页面:

  • 一开始用requests.get(urls)获取的r.content是页面首次加载的静态源码,循环里反复解析这个内容,自然只会抓到第一页标题。
  • 点击「Show more」后页面已经动态加载了新内容,但你没有重新获取当前浏览器的页面源码,还是用最开始的旧内容解析。

修正后的代码

from bs4 import BeautifulSoup
import pandas as pd
import time

# Selenium Routine
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.keys import Keys

# Removes SSL Issues With Chrome
options = webdriver.ChromeOptions()
options.add_argument('--ignore-certificate-errors')
options.add_argument('--ignore-ssl-errors')
options.add_argument('--ignore-certificate-errors-spki-list')
options.add_argument('log-level=3') 
options.add_argument('--disable-notifications')
#options.add_argument('--headless') # Comment to view browser actions

# 初始化浏览器并打开目标页面
urls = "https://jooble.org/blog/"
driver = webdriver.Chrome(executable_path="C:\webdrivers\chromedriver.exe", options=options)
driver.get(urls)

productlist = []

for i in range(1, 3):
    # 关键:每次循环都获取当前浏览器的页面源码,而不是初始静态内容
    soup = BeautifulSoup(driver.page_source, features='lxml')
    items = soup.find_all('div', class_='post')
    print(f'LOOP: start [{len(items)}]')

    for single_item in items:
        title = single_item.find('div', class_='front__news-title').text.strip()
        print('Title:', title)
        productlist.append({'Title': title})

    print()
    # 避免重复处理最后一页(如果循环到最后一次就不用点击了)
    if i < 2:
        time.sleep(2)  # 可以缩短等待时间,不需要5秒
        WebDriverWait(driver, 40).until(
            EC.element_to_be_clickable((By.XPATH, "//button[normalize-space()='Show more']"))
        ).send_keys(Keys.ENTER)

driver.close()

# Save Results
df = pd.DataFrame(productlist)
df.to_csv('Results.csv', index=False)

关键改动说明

  • 删掉了无用的requests.get(urls)和r变量,直接用driver.page_source获取当前浏览器的动态页面源码。
  • 每次循环都重新生成BeautifulSoup对象,确保解析的是最新加载的页面内容。
  • 添加了判断逻辑,最后一次循环不再点击「Show more」,避免不必要的等待。
  • 缩短了等待时间(从5秒改为2秒),提升效率,同时保留显式等待确保按钮可点击。

内容的提问来源于stack exchange,提问作者MarkWP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 04:30:52