You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫urlopen循环调用无效:重复爬取初始页面问题

为什么你的爬虫反复爬取初始页面?问题分析与修复方案

Hey there! Let's break down why your crawler is stuck scraping the first page over and over—it's not a problem with using urlopen in a for loop, but two key issues: missing request headers and flawed loop logic.

1. The Big Culprit: Missing Request Headers

Most modern websites check the User-Agent header to block non-browser requests. Your code uses uReq(Url) directly without setting any headers, which makes the site treat your request as suspicious. In many cases, sites will just return the first page content no matter what URL you request when they detect a missing or invalid User-Agent.

Fix for Headers

Wrap your request in a Request object and add a valid browser-like User-Agent to mimic a real visitor:

# Replace your original UClient = uReq(Url) lines with this
req = Request(Url, headers={
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
})
UClient = uReq(req)

2. Messy Loop Logic That's Sabotaging Pagination

Your code has redundant and confusing variable handling that's messing up the pagination flow:

  • The Url = URL_Next line in the else block does nothing—on the next loop iteration, you reassign Url using the if i == 0 check anyway.
  • The NumOfCrawledPages counter is unnecessary and makes your page number printing inaccurate.
  • When there's no next page, your loop still keeps running for the full 5 iterations instead of stopping early.

Refactored Loop Logic

Here's a cleaner version that fixes these issues:

import re
from math import ceil
from urllib.request import urlopen as uReq, Request
from bs4 import BeautifulSoup as soup

InitUrl = "https://mtgsingles.gr/search?q="
NumOfPages = 5
current_url = InitUrl

for page_num in range(NumOfPages):
    # Send request with proper headers
    req = Request(current_url, headers={
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    })
    UClient = uReq(req)
    page_html = UClient.read()
    UClient.close()

    page_soup = soup(page_html, "html.parser")
    cards = page_soup.findAll("div", {"class": ["iso-item", "item-row-view"]})

    # Extract card details (added basic error handling for robustness)
    for card in cards:
        try:
            card_name = card.div.div.strong.span.contents[3].contents[0].replace("\xa0 ", "")
        except (AttributeError, IndexError):
            card_name = "Unknown Card Name"
        
        try:
            if len(card.div.contents) > 3:
                cardP_T = card.div.contents[3].contents[1].text.replace("\n", "").strip()
            else:
                cardP_T = "Does not exist"
        except (AttributeError, IndexError):
            cardP_T = "Could not retrieve P/T"
        
        try:
            cardType = card.contents[3].text
        except (AttributeError, IndexError):
            cardType = "Could not retrieve Type"
        
        print(f"{card_name}\n{cardP_T}\n{cardType}\n")

    # Handle next page navigation
    try:
        next_link = page_soup.find("li", {"class": "next"}).a.get("href")
        current_url = f"https://mtgsingles.gr{next_link}"
        print(f"The next URL is: {current_url}\n")
    except AttributeError:
        print("Crawling process completed! No more information to retrieve!")
        break  # Stop loop early if no next page exists
    
    print(f"Moving to page : {page_num + 2}\n")  # Page numbers start at 1, so next is current +1

Bonus Tips for Robustness

  • Use find() instead of findAll() when you're looking for a single element (like the next page button)—it's faster and cleaner.
  • Add error handling to your card extraction code, like I did above. Web pages often have inconsistent structures, and this prevents your crawler from crashing unexpectedly.
  • Use f-strings for string formatting (Python 3.6+)—they're more readable than concatenation.

内容的提问来源于stack exchange,提问作者Petris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:19:03