You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Coursera数据分析课程爬虫问题:返回结果不稳定与索引越界报错

Hey there, let's break down what's going on with your web scraping code and fix those issues step by step!

First, why your original code had unstable results

The biggest culprit here is likely anti-scraping measures from Coursera. When you use urlopen without any request headers, your request looks like a bot instead of a real browser. Coursera might intentionally return incomplete or inconsistent page content in response, which explains why you got random pages of results each time you ran the code.

Second, why you hit "index out of range" errors

You noticed each course maps to 2 <h2> tags—so each page actually has 50 h2 elements (25 courses × 2 tags) instead of 25. Your hardcoded for x in range(26) loop was never a reliable way to fetch courses, and it would inevitably hit an index error when the page returned fewer h2s than expected (or more than you accounted for).

Fixed code with explanations

Let's rewrite the code to address both problems:

import requests
from bs4 import BeautifulSoup

courses_data_sci = []
# Use a session to persist connections, and add headers to mimic a real browser
session = requests.Session()
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# Loop through pages 1 to 9 (since you mentioned there are 9 total pages)
for page_num in range(1, 10):
    page_url = f"https://www.coursera.org/browse/data-science/data-analysis?languages=en&page={page_num}"
    try:
        # Fetch the page with headers to avoid being flagged as a bot
        response = session.get(page_url, headers=headers)
        response.raise_for_status()  # Catch HTTP errors like 403 or 500
        
        soup = BeautifulSoup(response.text, "html.parser")
        # Get all h2 tags under the rendered content area
        all_h2 = soup.select('#rendered-content h2')
        
        # Since each course has duplicate h2 tags, use a set to deduplicate course names
        unique_courses = {tag.get_text(strip=True) for tag in all_h2}
        courses_data_sci.extend(unique_courses)
        
        print(f'Page: {page_num} -- total courses collected so far: {len(courses_data_sci)}')
    except Exception as e:
        print(f'Failed to fetch page {page_num}: {str(e)}')

print(f'\nTotal unique courses collected: {len(courses_data_sci)}')

Key improvements:

  1. Added request headers: The User-Agent header makes your request look like it's coming from a Chrome browser, which should bypass Coursera's basic anti-scraping checks and return consistent page content.
  2. Deduplication with sets: Since each course has two identical h2 tags, using a set automatically removes duplicates so you only get each course name once.
  3. No hardcoded loops: Instead of guessing how many h2 tags to fetch, we directly iterate over all found tags—no more index errors.
  4. Error handling: The try/except block catches issues like failed requests, so you can see exactly which page caused problems.

Bonus: More precise targeting (optional)

If you want to avoid deduplication entirely, you can target the parent container of each course directly. Inspect the page in your browser's dev tools to find the class name for individual course cards (e.g., something like css-1rl8z0l—note that class names might change over time!). Then modify the parsing part:

# Replace the all_h2 and unique_courses lines with this:
course_cards = soup.select('#rendered-content .css-1rl8z0l')
for card in course_cards:
    course_name = card.find('h2').get_text(strip=True)
    courses_data_sci.append(course_name)

This way, you grab exactly one h2 per course card, no duplicates needed.

内容的提问来源于stack exchange,提问作者RichardS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:07:34