无分页滚动加载播客站点WebScraping全量数据获取方案咨询
爬取osimhistoria播客全部400条内容的解决方案
你的代码存在几个关键问题,导致只能获取前10条内容:
- 重复初始化Chrome驱动,覆盖了之前配置的无头模式,白白浪费了配置资源
- 未处理页面的滚动加载逻辑,完全没有触发"加载更多"按钮来获取后续内容
- 使用单个元素ID定位内容,只能拿到单一条目,无法批量获取所有播客
要爬完全部400条数据,核心逻辑是循环触发"加载更多"按钮,直到页面不再加载新内容,再统一提取所有条目信息。以下是修改后的完整代码:
import time import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.common.exceptions import NoSuchElementException # 配置Chrome无头模式,节省系统资源 options = Options() options.headless = True options.add_argument('window-size=1920x1080') options.add_argument('--disable-blink-features=AutomationControlled') # 规避网站自动化检测 website = "https://www.osimhistoria.com/osimhistoria" driver = webdriver.Chrome(options=options) driver.get(website) driver.maximize_window() # 存储播客数据的列表 podcast_data = [] # 循环点击加载更多按钮,直到按钮消失 while True: try: # 查找加载更多按钮,请根据页面实际按钮的文本/类名调整此XPath load_more_btn = driver.find_element(By.XPATH, '//button[contains(text(), "加载更多")]') # 滚动到按钮位置,确保按钮可被点击 driver.execute_script("arguments[0].scrollIntoView(true);", load_more_btn) time.sleep(2) # 等待滚动完成 load_more_btn.click() time.sleep(3) # 等待新内容加载渲染完成 except NoSuchElementException: # 没有加载更多按钮时,说明所有内容已加载完成,退出循环 break # 提取所有已加载的播客条目 # 这里使用你原代码中的容器class,实际条目为容器的子元素,请根据页面HTML结构调整定位 all_podcasts = driver.find_elements(By.CLASS_NAME, 'VM7gjN') for podcast in all_podcasts: # 提取所需信息,需根据页面实际元素结构调整定位规则 title = podcast.find_element(By.TAG_NAME, 'h3').text if podcast.find_elements(By.TAG_NAME, 'h3') else '无标题' publish_date = podcast.find_element(By.CLASS_NAME, '替换为实际日期类名').text if podcast.find_elements(By.CLASS_NAME, '替换为实际日期类名') else '无日期' audio_link = podcast.find_element(By.TAG_NAME, 'audio').get_attribute('src') if podcast.find_elements(By.TAG_NAME, 'audio') else '无音频链接' podcast_data.append({ '标题': title, '发布日期': publish_date, '音频链接': audio_link }) # 将数据保存为CSV文件 df = pd.DataFrame(podcast_data) df.to_csv('osimhistoria_podcasts.csv', index=False, encoding='utf-8-sig') print(f"共爬取到{len(podcast_data)}条播客数据,已保存至osimhistoria_podcasts.csv") driver.quit()
关键说明
- 元素定位调整:页面中的加载更多按钮、播客条目子元素(日期、音频链接)的选择器,需要你打开浏览器开发者工具,对照页面HTML结构修改,上述代码为通用示例
- 等待时间微调:如果网络速度较慢,可适当延长
time.sleep()的时长,避免内容未加载完成就开始提取 - 反爬规避:添加
--disable-blink-features=AutomationControlled参数,可降低被网站识别为自动化爬虫的概率
内容的提问来源于stack exchange,提问作者zachi
相关产品推荐
相关产品推荐

