You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无分页滚动加载播客站点WebScraping全量数据获取方案咨询

爬取osimhistoria播客全部400条内容的解决方案

你的代码存在几个关键问题,导致只能获取前10条内容:

  • 重复初始化Chrome驱动,覆盖了之前配置的无头模式,白白浪费了配置资源
  • 未处理页面的滚动加载逻辑,完全没有触发"加载更多"按钮来获取后续内容
  • 使用单个元素ID定位内容,只能拿到单一条目,无法批量获取所有播客

要爬完全部400条数据,核心逻辑是循环触发"加载更多"按钮,直到页面不再加载新内容,再统一提取所有条目信息。以下是修改后的完整代码:

import time
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.common.exceptions import NoSuchElementException

# 配置Chrome无头模式,节省系统资源
options = Options()
options.headless = True
options.add_argument('window-size=1920x1080')
options.add_argument('--disable-blink-features=AutomationControlled')  # 规避网站自动化检测

website = "https://www.osimhistoria.com/osimhistoria"
driver = webdriver.Chrome(options=options)
driver.get(website)
driver.maximize_window()

# 存储播客数据的列表
podcast_data = []

# 循环点击加载更多按钮,直到按钮消失
while True:
    try:
        # 查找加载更多按钮,请根据页面实际按钮的文本/类名调整此XPath
        load_more_btn = driver.find_element(By.XPATH, '//button[contains(text(), "加载更多")]')
        # 滚动到按钮位置,确保按钮可被点击
        driver.execute_script("arguments[0].scrollIntoView(true);", load_more_btn)
        time.sleep(2)  # 等待滚动完成
        load_more_btn.click()
        time.sleep(3)  # 等待新内容加载渲染完成
    except NoSuchElementException:
        # 没有加载更多按钮时,说明所有内容已加载完成,退出循环
        break

# 提取所有已加载的播客条目
# 这里使用你原代码中的容器class,实际条目为容器的子元素,请根据页面HTML结构调整定位
all_podcasts = driver.find_elements(By.CLASS_NAME, 'VM7gjN')

for podcast in all_podcasts:
    # 提取所需信息,需根据页面实际元素结构调整定位规则
    title = podcast.find_element(By.TAG_NAME, 'h3').text if podcast.find_elements(By.TAG_NAME, 'h3') else '无标题'
    publish_date = podcast.find_element(By.CLASS_NAME, '替换为实际日期类名').text if podcast.find_elements(By.CLASS_NAME, '替换为实际日期类名') else '无日期'
    audio_link = podcast.find_element(By.TAG_NAME, 'audio').get_attribute('src') if podcast.find_elements(By.TAG_NAME, 'audio') else '无音频链接'
    
    podcast_data.append({
        '标题': title,
        '发布日期': publish_date,
        '音频链接': audio_link
    })

# 将数据保存为CSV文件
df = pd.DataFrame(podcast_data)
df.to_csv('osimhistoria_podcasts.csv', index=False, encoding='utf-8-sig')

print(f"共爬取到{len(podcast_data)}条播客数据,已保存至osimhistoria_podcasts.csv")

driver.quit()

关键说明

  1. 元素定位调整:页面中的加载更多按钮、播客条目子元素(日期、音频链接)的选择器,需要你打开浏览器开发者工具,对照页面HTML结构修改,上述代码为通用示例
  2. 等待时间微调:如果网络速度较慢,可适当延长time.sleep()的时长,避免内容未加载完成就开始提取
  3. 反爬规避:添加--disable-blink-features=AutomationControlled参数,可降低被网站识别为自动化爬虫的概率

内容的提问来源于stack exchange,提问作者zachi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 18:01:21