Python爬虫urlopen循环调用无效:重复爬取初始页面问题
Hey there! Let's break down why your crawler is stuck scraping the first page over and over—it's not a problem with using urlopen in a for loop, but two key issues: missing request headers and flawed loop logic.
1. The Big Culprit: Missing Request Headers
Most modern websites check the User-Agent header to block non-browser requests. Your code uses uReq(Url) directly without setting any headers, which makes the site treat your request as suspicious. In many cases, sites will just return the first page content no matter what URL you request when they detect a missing or invalid User-Agent.
Fix for Headers
Wrap your request in a Request object and add a valid browser-like User-Agent to mimic a real visitor:
# Replace your original UClient = uReq(Url) lines with this req = Request(Url, headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' }) UClient = uReq(req)
2. Messy Loop Logic That's Sabotaging Pagination
Your code has redundant and confusing variable handling that's messing up the pagination flow:
- The
Url = URL_Nextline in theelseblock does nothing—on the next loop iteration, you reassignUrlusing theif i == 0check anyway. - The
NumOfCrawledPagescounter is unnecessary and makes your page number printing inaccurate. - When there's no next page, your loop still keeps running for the full 5 iterations instead of stopping early.
Refactored Loop Logic
Here's a cleaner version that fixes these issues:
import re from math import ceil from urllib.request import urlopen as uReq, Request from bs4 import BeautifulSoup as soup InitUrl = "https://mtgsingles.gr/search?q=" NumOfPages = 5 current_url = InitUrl for page_num in range(NumOfPages): # Send request with proper headers req = Request(current_url, headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' }) UClient = uReq(req) page_html = UClient.read() UClient.close() page_soup = soup(page_html, "html.parser") cards = page_soup.findAll("div", {"class": ["iso-item", "item-row-view"]}) # Extract card details (added basic error handling for robustness) for card in cards: try: card_name = card.div.div.strong.span.contents[3].contents[0].replace("\xa0 ", "") except (AttributeError, IndexError): card_name = "Unknown Card Name" try: if len(card.div.contents) > 3: cardP_T = card.div.contents[3].contents[1].text.replace("\n", "").strip() else: cardP_T = "Does not exist" except (AttributeError, IndexError): cardP_T = "Could not retrieve P/T" try: cardType = card.contents[3].text except (AttributeError, IndexError): cardType = "Could not retrieve Type" print(f"{card_name}\n{cardP_T}\n{cardType}\n") # Handle next page navigation try: next_link = page_soup.find("li", {"class": "next"}).a.get("href") current_url = f"https://mtgsingles.gr{next_link}" print(f"The next URL is: {current_url}\n") except AttributeError: print("Crawling process completed! No more information to retrieve!") break # Stop loop early if no next page exists print(f"Moving to page : {page_num + 2}\n") # Page numbers start at 1, so next is current +1
Bonus Tips for Robustness
- Use
find()instead offindAll()when you're looking for a single element (like the next page button)—it's faster and cleaner. - Add error handling to your card extraction code, like I did above. Web pages often have inconsistent structures, and this prevents your crawler from crashing unexpectedly.
- Use f-strings for string formatting (Python 3.6+)—they're more readable than concatenation.
内容的提问来源于stack exchange,提问作者Petris

