You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyCharm爬虫脚本分页功能失效求助

解决datacenters.com爬虫分页失效问题

问题根源

  1. 静态请求无法模拟JS点击:你用next_page_button.click()属于浏览器端的JS交互,requests是纯HTTP请求库,根本执行不了这个操作。
  2. 分页元素定位错误:原代码找的下一页按钮类名不准确,而且这个网站的分页是通过URL参数page实现的(比如第二页是.../virginia?page=2),不是靠按钮里的a标签跳转。
  3. 重复请求浪费资源:循环里重复对同一页面发请求,完全可以复用get_data_from_page里的soup对象。

修正后的代码

import requests
from bs4 import BeautifulSoup
import csv

# Base URL of the website
base_url = 'https://www.datacenters.com'
# URL template with page parameter
main_url_template = f'{base_url}/locations/united-states/virginia?page={{}}'

# 模拟浏览器请求头,避免被反爬
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

def get_data_from_page(url, writer):
    response = requests.get(url, headers=headers)
    if response.status_code == 200:
        soup = BeautifulSoup(response.content, 'html.parser')
        location_tiles = soup.find_all('div', class_='LocationTile__details__sXkB0')
        
        if not location_tiles:
            print(f"No location tiles found on {url}")
            return False  # 没有数据,返回False终止分页
        
        for tile in location_tiles:
            name = tile.find('div', class_='LocationTile__name__NrDKr').text.strip()
            address = tile.find('div', class_='LocationTile__address__Utj30').text.strip()
            
            parent_anchor = tile.find_parent('a', href=True)
            if parent_anchor:
                relative_link = parent_anchor['href']
                link = f'{base_url}{relative_link}'
                
                detail_response = requests.get(link, headers=headers)
                if detail_response.status_code == 200:
                    detail_soup = BeautifulSoup(detail_response.content, 'html.parser')
                    
                    power_div = detail_soup.find('div', id='power')
                    power = power_div.find('strong').text.strip() if power_div else 'N/A'
                    
                    sqf_div = detail_soup.find('div', id='statInfo')
                    sqf = sqf_div.find('strong').text.strip() if sqf_div else 'N/A'
                    
                    writer.writerow([name, address, power, sqf, link])
                    print(f"Scraped data for {name}")
                else:
                    print(f"Failed to retrieve details from {link}. Status code: {detail_response.status_code}")
            else:
                print(f"No link found for {name}")
        return True  # 有数据,继续分页
    else:
        print(f"Failed to retrieve the webpage. Status code: {response.status_code}")
        return False

# Open a CSV file to write the data
with open('datacenters.csv', mode='w', newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(['Name', 'Address', 'Power', 'SQF', 'Link'])
    
    page_num = 1
    while True:
        page_url = main_url_template.format(page_num)
        print(f"Scraping page: {page_url}")
        has_data = get_data_from_page(page_url, writer)
        if not has_data:
            print("No more pages or no data on current page. Stopping...")
            break
        page_num += 1

关键修改说明

  • 改用URL参数分页:直接通过?page=X的格式构造分页URL,避免依赖JS交互。
  • 添加请求头:模拟浏览器请求,降低被网站反爬拦截的概率。
  • 返回分页状态:让get_data_from_page返回是否还有数据,以此作为终止循环的条件。
  • 优化数据处理:对文本内容用strip()清理多余空格,避免CSV里的脏数据。

内容的提问来源于stack exchange,提问作者Cooper Scranton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 22:47:09