爬取苹果应用商店数据时出现No connection adapters were found错误求助
Hey there, let's break down why you're running into this connection error and fix it step by step!
The Root Cause
Looking at your code, I spotted the key issue: you're trying to replace the itms-appss protocol AFTER making the requests.get() call. When you run r = requests.get(line[0].strip(),headers=headers).text, you're using the original (broken) URL from your CSV file—your replacement logic runs too late and never actually affects the request you're sending. On top of that, your URL has & (an HTML-encoded ampersand) instead of &, which can also confuse the requests library.
Step-by-Step Fixes
Clean URLs Before Sending Requests
Fix the protocol and encoded characters first, then use the cleaned URL for your request.Improve Error Handling
A bareexcept Exception: passhides all issues, making debugging impossible. Instead, catch specific errors and log details like which URL failed.Add Rate Limiting
Apple's servers may block rapid repeated requests. Adding a small delay helps avoid being flagged as a bot.
Modified Working Code
Here's your updated code with all the fixes:
import csv import requests from bs4 import BeautifulSoup import time with open('App_Store_Links.csv', newline='') as f_urls, open('appsinfo.csv', 'w', newline='') as f_output: csv_urls = csv.reader(f_urls) csv_output = csv.writer(f_output) csv_output.writerow(['App Name', 'Category','Size','Developer','Age Rating','Rating','Rating Numbers']) headers = requests.utils.default_headers() headers['User-Agent'] = 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.87 Safari/537.36' # Skip header row if your CSV has one (uncomment this line if needed) # next(csv_urls) for line in csv_urls: # Clean up the URL FIRST before making the request raw_url = line[0].strip() cleaned_url = raw_url.replace("itms-appss", "https").replace("&", "&") try: # Send request with the cleaned, valid HTTPS URL r = requests.get(cleaned_url, headers=headers) # Check if the request was successful (catches 404/403 errors) r.raise_for_status() soup = BeautifulSoup(r.text, 'lxml') # Extract app info with safety checks to avoid index errors info_items = soup.findAll('dd', class_='information-list__item__definition l-column medium-9 large-6') app_name = soup.find('h1', class_='product-header__title app-header__title').text.strip() if soup.find('h1', class_='product-header__title app-header__title') else "N/A" Category = info_items[2].text.strip() if len(info_items) > 2 else "N/A" Size = info_items[1].text.strip() if len(info_items) > 1 else "N/A" Developer = info_items[0].text.strip() if len(info_items) > 0 else "N/A" Age_Rating = soup.find('span', class_='badge badge--product-title').text.strip() if soup.find('span', class_='badge badge--product-title') else "N/A" Rating = soup.find('span', class_='we-customer-ratings__averages__display').text.strip() if soup.find('span', class_='we-customer-ratings__averages__display') else "N/A" Rating_number = soup.find('div', class_='we-customer-ratings__count small-hide medium-show').text.strip() if soup.find('div', class_='we-customer-ratings__count small-hide medium-show') else "N/A" csv_output.writerow([app_name, Category, Size, Developer, Age_Rating, Rating, Rating_number]) # Add a 2-second delay to avoid triggering anti-bot measures time.sleep(2) except requests.exceptions.RequestException as e: print(f"Failed to fetch {cleaned_url}: {str(e)}") # Log errors to your output CSV for later review csv_output.writerow(["ERROR", "ERROR", "ERROR", "ERROR", "ERROR", "ERROR", f"Request failed: {str(e)}"]) except Exception as e: print(f"Error parsing {cleaned_url}: {str(e)}") csv_output.writerow(["ERROR", "ERROR", "ERROR", "ERROR", "ERROR", "ERROR", f"Parsing failed: {str(e)}"])
Key Improvements Explained
- URL Cleaning First: We fix the broken protocol and encoded ampersand before sending the request, so
requestsgets a valid HTTPS URL it can process. - Request Validation:
r.raise_for_status()alerts you to server errors like 404 (missing page) or 403 (forbidden), which helps you spot broken links or blocks. - Graceful Info Extraction: We add checks for each element (e.g.,
if len(info_items) > 2) to avoid crashing if an app page is missing certain details. - Rate Limiting:
time.sleep(2)slows down requests to avoid being blocked by Apple's anti-bot systems. - Useful Error Logging: Instead of ignoring errors, we print details and write error rows to your CSV, so you can follow up on problematic URLs.
Additional Tips
- If you still get blocked, try rotating
User-Agentstrings (use a list of common user agents and pick one at random for each request). - For large-scale scraping, consider using a proxy service to avoid IP bans.
- Always check Apple's robots.txt to ensure you're respecting their scraping policies.
内容的提问来源于stack exchange,提问作者Wisam Barakat

