使用Python爬取网页表格:获取泰晤士2018完整大学排名表
Hey there, I’ve worked through how to pull the full set of 2018 THE rankings (over 1000 schools) using the tools you mentioned. The key here is handling the site’s pagination system (25 schools per page) and making sure our requests don’t get blocked by basic anti-scraping measures. Here’s a complete, working solution:
Step 1: Import Required Libraries
First, we’ll need requests for fetching pages, BeautifulSoup for parsing HTML, time to add delays between requests, and csv if you want to save results to a file.
import requests from bs4 import BeautifulSoup import time import csv
Step 2: Configure Request Headers
Many sites block requests without a proper User-Agent header, so we’ll add one to mimic a real browser:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' }
Step 3: Scrape All Pages
We’ll loop through each page starting from page 0, fetch content, parse table rows, and stop when we hit a page with no ranking data (meaning we’ve reached the end).
base_url = "https://www.timeshighereducation.com/world-university-rankings/2018/world-ranking#!/page/{page}/length/25/sort_by/rank/sort_order/asc/cols/scores" all_rankings = [] page = 0 while True: # Fetch the target page url = base_url.format(page=page) response = requests.get(url, headers=headers) # Check if request succeeded if response.status_code != 200: print(f"Failed to fetch page {page}: Status code {response.status_code}") break soup = BeautifulSoup(response.content, 'html.parser') # Grab table rows (odd/even classes alternate rows) table_rows = soup.select('tr.odd, tr.even') if not table_rows: print(f"No more data found at page {page}. Stopping.") break # Extract data from each row for row in table_rows: ranking_data = {} # Extract rank (handles range ranks like "101-125") rank_cell = row.select_one('td.rank') ranking_data['rank'] = rank_cell.get_text(strip=True) if rank_cell else "N/A" # Extract university name and link name_cell = row.select_one('td.institution a') ranking_data['university_name'] = name_cell.get_text(strip=True) if name_cell else "N/A" ranking_data['university_link'] = name_cell['href'] if name_cell else "N/A" # Extract country country_cell = row.select_one('td.country') ranking_data['country'] = country_cell.get_text(strip=True) if country_cell else "N/A" # Extract all score metrics score_cells = row.select('td.score') if len(score_cells) >= 5: ranking_data['teaching_score'] = score_cells[0].get_text(strip=True) ranking_data['research_score'] = score_cells[1].get_text(strip=True) ranking_data['citations_score'] = score_cells[2].get_text(strip=True) ranking_data['industry_income_score'] = score_cells[3].get_text(strip=True) ranking_data['international_outlook_score'] = score_cells[4].get_text(strip=True) all_rankings.append(ranking_data) print(f"Successfully scraped page {page} ({len(table_rows)} schools)") page += 1 # Add a small delay to avoid overwhelming the server time.sleep(2)
Step 4: Save Results to CSV (Optional)
Once scraping is done, save the data to a CSV file for easy analysis:
# Define CSV column order csv_columns = ['rank', 'university_name', 'country', 'teaching_score', 'research_score', 'citations_score', 'industry_income_score', 'international_outlook_score', 'university_link'] with open('2018_the_world_university_rankings.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=csv_columns) writer.writeheader() for data in all_rankings: writer.writerow(data) print(f"Scraping complete! Saved {len(all_rankings)} schools to '2018_the_world_university_rankings.csv'")
Key Notes:
- Anti-Scraping: The
time.sleep(2)helps prevent your IP from being blocked—adjust it if needed, but don’t make it too short. - Range Ranks: The code captures ranks like "101-125" as-is, since that’s how they’re displayed on the site.
- Error Handling: Missing cells default to "N/A" to avoid crashes if the page structure has inconsistencies.
内容的提问来源于stack exchange,提问作者Alexander Kalinovskiy

