如何用Python Playwright处理URL不变的网页分页爬取问题?
解决Playwright爬取单页应用分页问题
针对这个单页应用的分页场景,核心思路是模拟点击下一页按钮,循环提取每一页数据,直到没有下一页为止。以下是修改后的完整代码:
from bs4 import BeautifulSoup as bs from playwright.sync_api import sync_playwright url = 'https://franchisedisclosure.gov.au/Register' # 初始化存储数据的列表 names = [] industries = [] locations = [] with sync_playwright() as p: browser = p.chromium.launch(headless=False, slow_mo=50) page = browser.new_page() page.goto(url) # 同意条款并进入页面 page.locator("text=I agree to the terms of use").click() page.locator("text=Continue").click() page.wait_for_load_state('domcontentloaded') page.wait_for_selector('tbody') # 等待表格内容加载完成 while True: # 提取当前页的表格数据 html = page.inner_html('table.table.table-hover') soup = bs(html, 'html.parser') table = soup.find('tbody') rows = table.findAll('tr') for row in rows: info = row.findAll('td') names.append(info[0].text.strip()) industries.append(info[1].text.strip()) locations.append(info[2].text.strip()) # 定位下一页按钮,判断是否可点击 next_button = page.locator('ul.pagination li.page-item:last-child a.page-link') # 检查按钮是否被禁用(通过父元素的disabled类判断) is_disabled = next_button.locator('..').get_attribute('class').then(lambda cls: 'disabled' in cls) if is_disabled: break # 没有下一页,终止循环 # 点击下一页,等待表格内容更新 next_button.click() page.wait_for_selector('tbody', state='attached') # 等待新页面数据加载 # 加短延迟确保数据完全渲染,避免提取到旧数据 page.wait_for_timeout(500) browser.close() # 打印结果示例 print(f"共爬取{len(names)}条数据") print("第一条数据:", names[0], industries[0], locations[0])
关键说明:
- 变量修复:原代码中
industry和Locations被循环赋值覆盖,改为append到列表中,确保所有数据都被存储。 - 分页判断逻辑:通过检查下一页按钮的父元素是否包含
disabled类,判断是否还有下一页。如果按钮不可点击,终止循环。 - 等待机制:每次点击下一页后,等待表格元素重新加载,并添加短延迟确保数据完全渲染,避免提取到旧页面的数据。
- 循环提取:用
while True循环处理每一页,直到没有下一页为止。
内容的提问来源于stack exchange,提问作者Chiedozie Agwu
相关产品推荐
相关产品推荐

