PyCharm爬虫脚本分页功能失效求助
解决datacenters.com爬虫分页失效问题
问题根源
- 静态请求无法模拟JS点击:你用
next_page_button.click()属于浏览器端的JS交互,requests是纯HTTP请求库,根本执行不了这个操作。 - 分页元素定位错误:原代码找的下一页按钮类名不准确,而且这个网站的分页是通过URL参数
page实现的(比如第二页是.../virginia?page=2),不是靠按钮里的a标签跳转。 - 重复请求浪费资源:循环里重复对同一页面发请求,完全可以复用
get_data_from_page里的soup对象。
修正后的代码
import requests from bs4 import BeautifulSoup import csv # Base URL of the website base_url = 'https://www.datacenters.com' # URL template with page parameter main_url_template = f'{base_url}/locations/united-states/virginia?page={{}}' # 模拟浏览器请求头,避免被反爬 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } def get_data_from_page(url, writer): response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') location_tiles = soup.find_all('div', class_='LocationTile__details__sXkB0') if not location_tiles: print(f"No location tiles found on {url}") return False # 没有数据,返回False终止分页 for tile in location_tiles: name = tile.find('div', class_='LocationTile__name__NrDKr').text.strip() address = tile.find('div', class_='LocationTile__address__Utj30').text.strip() parent_anchor = tile.find_parent('a', href=True) if parent_anchor: relative_link = parent_anchor['href'] link = f'{base_url}{relative_link}' detail_response = requests.get(link, headers=headers) if detail_response.status_code == 200: detail_soup = BeautifulSoup(detail_response.content, 'html.parser') power_div = detail_soup.find('div', id='power') power = power_div.find('strong').text.strip() if power_div else 'N/A' sqf_div = detail_soup.find('div', id='statInfo') sqf = sqf_div.find('strong').text.strip() if sqf_div else 'N/A' writer.writerow([name, address, power, sqf, link]) print(f"Scraped data for {name}") else: print(f"Failed to retrieve details from {link}. Status code: {detail_response.status_code}") else: print(f"No link found for {name}") return True # 有数据,继续分页 else: print(f"Failed to retrieve the webpage. Status code: {response.status_code}") return False # Open a CSV file to write the data with open('datacenters.csv', mode='w', newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(['Name', 'Address', 'Power', 'SQF', 'Link']) page_num = 1 while True: page_url = main_url_template.format(page_num) print(f"Scraping page: {page_url}") has_data = get_data_from_page(page_url, writer) if not has_data: print("No more pages or no data on current page. Stopping...") break page_num += 1
关键修改说明
- 改用URL参数分页:直接通过
?page=X的格式构造分页URL,避免依赖JS交互。 - 添加请求头:模拟浏览器请求,降低被网站反爬拦截的概率。
- 返回分页状态:让
get_data_from_page返回是否还有数据,以此作为终止循环的条件。 - 优化数据处理:对文本内容用
strip()清理多余空格,避免CSV里的脏数据。
内容的提问来源于stack exchange,提问作者Cooper Scranton
相关产品推荐
相关产品推荐

