You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取网页表格:获取泰晤士2018完整大学排名表

Scraping the 2018 THE World University Rankings (1000+ Universities) with Requests & BeautifulSoup

Hey there, I’ve worked through how to pull the full set of 2018 THE rankings (over 1000 schools) using the tools you mentioned. The key here is handling the site’s pagination system (25 schools per page) and making sure our requests don’t get blocked by basic anti-scraping measures. Here’s a complete, working solution:

Step 1: Import Required Libraries

First, we’ll need requests for fetching pages, BeautifulSoup for parsing HTML, time to add delays between requests, and csv if you want to save results to a file.

import requests
from bs4 import BeautifulSoup
import time
import csv

Step 2: Configure Request Headers

Many sites block requests without a proper User-Agent header, so we’ll add one to mimic a real browser:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}

Step 3: Scrape All Pages

We’ll loop through each page starting from page 0, fetch content, parse table rows, and stop when we hit a page with no ranking data (meaning we’ve reached the end).

base_url = "https://www.timeshighereducation.com/world-university-rankings/2018/world-ranking#!/page/{page}/length/25/sort_by/rank/sort_order/asc/cols/scores"
all_rankings = []

page = 0
while True:
    # Fetch the target page
    url = base_url.format(page=page)
    response = requests.get(url, headers=headers)
    
    # Check if request succeeded
    if response.status_code != 200:
        print(f"Failed to fetch page {page}: Status code {response.status_code}")
        break
    
    soup = BeautifulSoup(response.content, 'html.parser')
    
    # Grab table rows (odd/even classes alternate rows)
    table_rows = soup.select('tr.odd, tr.even')
    if not table_rows:
        print(f"No more data found at page {page}. Stopping.")
        break
    
    # Extract data from each row
    for row in table_rows:
        ranking_data = {}
        
        # Extract rank (handles range ranks like "101-125")
        rank_cell = row.select_one('td.rank')
        ranking_data['rank'] = rank_cell.get_text(strip=True) if rank_cell else "N/A"
        
        # Extract university name and link
        name_cell = row.select_one('td.institution a')
        ranking_data['university_name'] = name_cell.get_text(strip=True) if name_cell else "N/A"
        ranking_data['university_link'] = name_cell['href'] if name_cell else "N/A"
        
        # Extract country
        country_cell = row.select_one('td.country')
        ranking_data['country'] = country_cell.get_text(strip=True) if country_cell else "N/A"
        
        # Extract all score metrics
        score_cells = row.select('td.score')
        if len(score_cells) >= 5:
            ranking_data['teaching_score'] = score_cells[0].get_text(strip=True)
            ranking_data['research_score'] = score_cells[1].get_text(strip=True)
            ranking_data['citations_score'] = score_cells[2].get_text(strip=True)
            ranking_data['industry_income_score'] = score_cells[3].get_text(strip=True)
            ranking_data['international_outlook_score'] = score_cells[4].get_text(strip=True)
        
        all_rankings.append(ranking_data)
    
    print(f"Successfully scraped page {page} ({len(table_rows)} schools)")
    page += 1
    # Add a small delay to avoid overwhelming the server
    time.sleep(2)

Step 4: Save Results to CSV (Optional)

Once scraping is done, save the data to a CSV file for easy analysis:

# Define CSV column order
csv_columns = ['rank', 'university_name', 'country', 'teaching_score', 'research_score', 'citations_score', 'industry_income_score', 'international_outlook_score', 'university_link']

with open('2018_the_world_university_rankings.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.DictWriter(csvfile, fieldnames=csv_columns)
    writer.writeheader()
    for data in all_rankings:
        writer.writerow(data)

print(f"Scraping complete! Saved {len(all_rankings)} schools to '2018_the_world_university_rankings.csv'")

Key Notes:

  • Anti-Scraping: The time.sleep(2) helps prevent your IP from being blocked—adjust it if needed, but don’t make it too short.
  • Range Ranks: The code captures ranks like "101-125" as-is, since that’s how they’re displayed on the site.
  • Error Handling: Missing cells default to "N/A" to avoid crashes if the page structure has inconsistencies.

内容的提问来源于stack exchange,提问作者Alexander Kalinovskiy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:05:12