You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取苹果应用商店数据时出现No connection adapters were found错误求助

Fixing the "No connection adapters were found" Error in App Store Scraping

Hey there, let's break down why you're running into this connection error and fix it step by step!

The Root Cause

Looking at your code, I spotted the key issue: you're trying to replace the itms-appss protocol AFTER making the requests.get() call. When you run r = requests.get(line[0].strip(),headers=headers).text, you're using the original (broken) URL from your CSV file—your replacement logic runs too late and never actually affects the request you're sending. On top of that, your URL has & (an HTML-encoded ampersand) instead of &, which can also confuse the requests library.

Step-by-Step Fixes

  1. Clean URLs Before Sending Requests
    Fix the protocol and encoded characters first, then use the cleaned URL for your request.

  2. Improve Error Handling
    A bare except Exception: pass hides all issues, making debugging impossible. Instead, catch specific errors and log details like which URL failed.

  3. Add Rate Limiting
    Apple's servers may block rapid repeated requests. Adding a small delay helps avoid being flagged as a bot.

Modified Working Code

Here's your updated code with all the fixes:

import csv
import requests
from bs4 import BeautifulSoup
import time

with open('App_Store_Links.csv', newline='') as f_urls, open('appsinfo.csv', 'w', newline='') as f_output: 
    csv_urls = csv.reader(f_urls)
    csv_output = csv.writer(f_output)
    csv_output.writerow(['App Name', 'Category','Size','Developer','Age Rating','Rating','Rating Numbers'])
    
    headers = requests.utils.default_headers()
    headers['User-Agent'] = 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.87 Safari/537.36'
    
    # Skip header row if your CSV has one (uncomment this line if needed)
    # next(csv_urls)
    
    for line in csv_urls:
        # Clean up the URL FIRST before making the request
        raw_url = line[0].strip()
        cleaned_url = raw_url.replace("itms-appss", "https").replace("&", "&")
        
        try:
            # Send request with the cleaned, valid HTTPS URL
            r = requests.get(cleaned_url, headers=headers)
            # Check if the request was successful (catches 404/403 errors)
            r.raise_for_status()
            soup = BeautifulSoup(r.text, 'lxml')
            
            # Extract app info with safety checks to avoid index errors
            info_items = soup.findAll('dd', class_='information-list__item__definition l-column medium-9 large-6')
            app_name = soup.find('h1', class_='product-header__title app-header__title').text.strip() if soup.find('h1', class_='product-header__title app-header__title') else "N/A"
            Category = info_items[2].text.strip() if len(info_items) > 2 else "N/A"
            Size = info_items[1].text.strip() if len(info_items) > 1 else "N/A"
            Developer = info_items[0].text.strip() if len(info_items) > 0 else "N/A"
            Age_Rating = soup.find('span', class_='badge badge--product-title').text.strip() if soup.find('span', class_='badge badge--product-title') else "N/A"
            Rating = soup.find('span', class_='we-customer-ratings__averages__display').text.strip() if soup.find('span', class_='we-customer-ratings__averages__display') else "N/A"
            Rating_number = soup.find('div', class_='we-customer-ratings__count small-hide medium-show').text.strip() if soup.find('div', class_='we-customer-ratings__count small-hide medium-show') else "N/A"
            
            csv_output.writerow([app_name, Category, Size, Developer, Age_Rating, Rating, Rating_number])
            
            # Add a 2-second delay to avoid triggering anti-bot measures
            time.sleep(2)
            
        except requests.exceptions.RequestException as e:
            print(f"Failed to fetch {cleaned_url}: {str(e)}")
            # Log errors to your output CSV for later review
            csv_output.writerow(["ERROR", "ERROR", "ERROR", "ERROR", "ERROR", "ERROR", f"Request failed: {str(e)}"])
        except Exception as e:
            print(f"Error parsing {cleaned_url}: {str(e)}")
            csv_output.writerow(["ERROR", "ERROR", "ERROR", "ERROR", "ERROR", "ERROR", f"Parsing failed: {str(e)}"])

Key Improvements Explained

  • URL Cleaning First: We fix the broken protocol and encoded ampersand before sending the request, so requests gets a valid HTTPS URL it can process.
  • Request Validation: r.raise_for_status() alerts you to server errors like 404 (missing page) or 403 (forbidden), which helps you spot broken links or blocks.
  • Graceful Info Extraction: We add checks for each element (e.g., if len(info_items) > 2) to avoid crashing if an app page is missing certain details.
  • Rate Limiting: time.sleep(2) slows down requests to avoid being blocked by Apple's anti-bot systems.
  • Useful Error Logging: Instead of ignoring errors, we print details and write error rows to your CSV, so you can follow up on problematic URLs.

Additional Tips

  • If you still get blocked, try rotating User-Agent strings (use a list of common user agents and pick one at random for each request).
  • For large-scale scraping, consider using a proxy service to avoid IP bans.
  • Always check Apple's robots.txt to ensure you're respecting their scraping policies.

内容的提问来源于stack exchange,提问作者Wisam Barakat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:24:10