You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析内容时分页信息无法获取的问题

解决BeautifulSoup爬取分页内容失败的问题

你遇到的分页获取失败,核心原因大概率是分页内容通过JavaScript动态加载,而requests仅能获取页面初始HTML,无法执行JS渲染后续内容;另外你当前使用的选择器可能也未命中正确的分页元素。以下是两种可行的解决方案:


方案1:用Selenium渲染动态内容

Selenium可模拟浏览器行为,完整加载并执行页面JS,获取包含动态分页的完整页面内容。

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import pandas as pd

# 初始化Chrome浏览器(需提前下载对应版本的chromedriver)
driver = webdriver.Chrome()
base_url = "https://www.yellowpages.lk/Medical.php"
driver.get(base_url)
time.sleep(3)  # 等待页面初始加载

# 初始化数据存储列表
C_List = []
D_List = []
N_List = []
A_List = []
T_List = []
W_List = []
V_List = []
M_List = []

while True:
    # 获取当前页面的完整HTML
    soup = BeautifulSoup(driver.page_source, "lxml")
    # 解析表格数据
    table_rows = soup.find('table', {'id': 'MedicalFacility'}).find('tbody').find_all('tr')
    for row in table_rows:
        tds = row.find_all('td')
        if len(tds) < 8:
            continue  # 跳过无效行
        
        C_List.append(tds[0].get_text().strip())
        D_List.append(tds[1].get_text().strip())
        N_List.append(tds[2].get_text().strip())
        A_List.append(tds[3].get_text().strip())
        T_List.append(tds[4].get_text().strip())
        W_List.append(tds[5].get_text().strip())
        V_List.append(tds[6].get_text().strip())
        M_List.append(tds[7].get_text().strip())
    
    # 尝试点击下一页
    try:
        # 等待下一页按钮可点击(选择器需根据实际页面调整)
        next_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, "li.page-item.next a"))
        )
        # 判断是否为最后一页
        if 'disabled' in next_btn.parent.get_attribute('class'):
            break
        next_btn.click()
        time.sleep(2)  # 等待翻页后页面加载
    except Exception as e:
        print("已到最后一页或无法定位下一页按钮:", e)
        break

# 关闭浏览器并保存数据
driver.quit()
df = pd.DataFrame({
    'Category': C_List,
    'District': D_List,
    'Name': N_List,
    'Address': A_List,
    'Telephone': T_List,
    'Whatsapp': W_List,
    'Viber': V_List,
    'MoH_Division': M_List
})
print(df)

方案2:抓包获取分页接口(效率更高)

通过浏览器开发者工具(F12)的Network面板,观察翻页时的XHR请求,找到直接返回数据的分页接口,用requests直接请求接口数据。

假设分页参数为page,示例代码如下:

from bs4 import BeautifulSoup
import requests
import pandas as pd
import time

base_url = "https://www.yellowpages.lk/Medical.php"
page_num = 1
# 初始化数据存储列表
C_List = []
D_List = []
N_List = []
A_List = []
T_List = []
W_List = []
V_List = []
M_List = []

while True:
    url = f"{base_url}?page={page_num}"
    # 添加请求头模拟浏览器,避免被拦截
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    response = requests.get(url, headers=headers)
    if response.status_code != 200:
        print(f"请求第{page_num}页失败,状态码:{response.status_code}")
        break
    
    soup = BeautifulSoup(response.content, "lxml")
    table_rows = soup.find('table', {'id': 'MedicalFacility'}).find('tbody').find_all('tr')
    if not table_rows:
        print("已到最后一页,无数据")
        break
    
    # 解析当前页数据
    for row in table_rows:
        tds = row.find_all('td')
        if len(tds) < 8:
            continue
        
        C_List.append(tds[0].get_text().strip())
        D_List.append(tds[1].get_text().strip())
        N_List.append(tds[2].get_text().strip())
        A_List.append(tds[3].get_text().strip())
        T_List.append(tds[4].get_text().strip())
        W_List.append(tds[5].get_text().strip())
        V_List.append(tds[6].get_text().strip())
        M_List.append(tds[7].get_text().strip())
    
    # 检查是否存在下一页
    pagination = soup.find('ul', class_='pagination')
    if not pagination or 'Next' not in pagination.get_text():
        break
    
    page_num += 1
    time.sleep(2)  # 控制请求频率,避免被封禁

# 保存数据
df = pd.DataFrame({
    'Category': C_List,
    'District': D_List,
    'Name': N_List,
    'Address': A_List,
    'Telephone': T_List,
    'Whatsapp': W_List,
    'Viber': V_List,
    'MoH_Division': M_List
})
print(df)

注意事项

  1. 使用Selenium时,需下载对应浏览器的驱动程序(如Chrome的chromedriver),并确保驱动版本与浏览器版本匹配。
  2. 两种方案都需添加合理的等待时间(time.sleep()),避免请求频率过高被网站封禁。
  3. 若页面分页元素结构变化,需及时调整CSS选择器或分页参数。

内容的提问来源于stack exchange,提问作者Prog_Beginner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 14:32:20