You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Selenium实现同一URL下多页面爬取

解决Selenium爬虫的翻页逻辑整合问题

问题说明

现有一段Python Selenium爬虫代码,可从指定网站抓取信息,但目标内容分布在61个页面中,页面URL固定,需通过点击“Next”按钮翻页。已知可用driver.find_element_by_link_text('Next').click()实现翻页,但不清楚如何将该逻辑整合到现有代码中。原代码如下:

from selenium import webdriver
from selenium.webdriver.support.wait import WebDriverWait    
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time

driver = webdriver.Chrome()
driver.get('https://mspotrace.org.my/Sccs_list')
time.sleep(20)

# Get list of elements
elements = WebDriverWait(driver, 20).until(EC.presence_of_all_elements_located((By.XPATH, "//a[@title='View on Map']")))

# Loop through element popups and pull details of facilities into DF
pos = 0
df = pd.DataFrame(columns=['facility_name','other_details'])

for element in elements:
    try: 
        data = []
        element.click()
        time.sleep(10)
        facility_name = driver.find_element_by_xpath('//h4[@class="modal-title"]').text
        other_details = driver.find_element_by_xpath('//div[@class="modal-body"]').text
        time.sleep(5)
        data.append(facility_name)
        data.append(other_details)
        df.loc[pos] = data
        WebDriverWait(driver,10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button[aria-label='Close'] > span"))).click() # close popup window
        print("Scraping info for",facility_name,"")
        time.sleep(15)
        pos+=1

    except Exception:
        alert = driver.switch_to.alert
        print("No geo location information")
        alert.accept()
        pass
       
print(df)

修改后的代码(整合翻页逻辑)

通过外层循环控制翻页次数,完成当前页抓取后点击“Next”进入下一页,直到遍历完61个页面:

from selenium import webdriver
from selenium.webdriver.support.wait import WebDriverWait    
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time

driver = webdriver.Chrome()
driver.get('https://mspotrace.org.my/Sccs_list')
# 用显式等待替代固定sleep,等待页面核心元素加载
WebDriverWait(driver, 20).until(EC.presence_of_element_located((By.XPATH, "//a[@title='View on Map']")))

df = pd.DataFrame(columns=['facility_name','other_details'])
total_pages = 61  # 目标总页数
current_page = 1

while current_page <= total_pages:
    print(f"正在抓取第 {current_page} 页内容...")
    # 获取当前页的所有目标元素
    elements = WebDriverWait(driver, 20).until(EC.presence_of_all_elements_located((By.XPATH, "//a[@title='View on Map']")))
    
    # 遍历当前页元素,抓取详情
    for element in elements:
        try: 
            data = []
            element.click()
            # 等待弹窗加载完成
            WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//h4[@class="modal-title"]')))
            
            facility_name = driver.find_element_by_xpath('//h4[@class="modal-title"]').text
            other_details = driver.find_element_by_xpath('//div[@class="modal-body"]').text
            
            data.append(facility_name)
            data.append(other_details)
            # 用append简化DataFrame数据添加
            df = df.append(pd.Series(data, index=df.columns), ignore_index=True)
            
            # 关闭弹窗
            close_btn = WebDriverWait(driver,10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button[aria-label='Close'] > span")))
            close_btn.click()
            print(f"已抓取:{facility_name}")
            time.sleep(2)  # 短暂等待避免操作过快

        except Exception as e:
            # 处理无地理位置的弹窗警告
            try:
                alert = driver.switch_to.alert
                print(f"{locals().get('facility_name', '未知设施')}:无地理位置信息")
                alert.accept()
            except:
                print(f"抓取出错:{str(e)}")
            continue
    
    # 翻页操作:未到最后一页时点击Next
    if current_page < total_pages:
        try:
            next_btn = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next')))
            next_btn.click()
            # 等待页面刷新,确认当前页元素失效
            WebDriverWait(driver, 20).until(EC.staleness_of(elements[0]))
            current_page += 1
            time.sleep(3)
        except Exception as e:
            print(f"翻页失败:{str(e)}")
            break
    else:
        break

# 保存结果到CSV(可选)
df.to_csv('facilities_info.csv', index=False, encoding='utf-8-sig')
print(f"抓取完成,共获取 {len(df)} 条数据")
driver.quit()

关键修改说明

  • 新增外层while循环控制总页数,确保遍历完所有目标页面
  • 用显式等待替代大部分固定time.sleep,提升代码稳定性和执行效率
  • 优化DataFrame数据添加方式,用append替代loc[pos]更简洁
  • 增加翻页异常处理,避免因按钮无法点击导致程序崩溃
  • 添加CSV保存逻辑,方便后续数据处理

内容的提问来源于stack exchange,提问作者Funkeh-Monkeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 03:18:23