爬虫脚本连接超时中断问题及解决方案咨询
Hey there, let's figure out why your scraping script keeps crashing and fix this properly. First off, I spotted a critical issue in your code that's undermining the retry setup you added—let's start there.
The Root Issue in Your Code
You set up a requests.Session with retry logic, but then you don't use it for the actual request that's failing! Look at this line:
soup = BeautifulSoup(requests.get(url).content, 'html.parser')
This calls requests.get() directly, bypassing your session entirely. That means none of your retry rules apply here—so when this request hits a timeout, it fails immediately without retrying. No wonder you're still seeing those errors!
Fixed Scraping Function
Let's rewrite your scrape function to use the session consistently, add timeouts, and include error handling to prevent the whole script from crashing:
def scrape(page): session = requests.Session() retry = Retry( connect=5, backoff_factor=0.5, status_forcelist=[500, 502, 503, 504] # Also retry on common server errors ) adapter = HTTPAdapter(max_retries=retry) session.mount('http://', adapter) session.mount('https://', adapter) try: # Use the session for ALL requests, and add explicit timeouts response = session.get(page, timeout=(10, 30)) # 10s connect timeout, 30s read timeout response.raise_for_status() # Trigger an error for 4xx/5xx status codes soup = BeautifulSoup(response.content, 'html.parser') return soup # Return your parsed soup or data object here except requests.exceptions.RequestException as e: # Log the error instead of crashing the whole script print(f"Failed to scrape {page}: {str(e)}") return None # Return a placeholder to indicate failure
Additional Fixes to Prevent Crashes
Even with the fixed function, scraping 8000 URLs can hit network issues or server limits. Here are more steps to make your script robust:
Add Rate Limiting: Target servers will block you if you hit them too fast. Add a small delay between requests to avoid overwhelming the server:
import time # After processing each URL in your CSV loop time.sleep(0.75) # Adjust between 0.5-2 seconds based on the site's toleranceTrack Failed URLs: Don't lose track of which URLs failed—save them to a file so you can retry later:
failed_urls = [] # Inside your CSV processing loop: soup = scrape(row[1]) if soup is None: failed_urls.append(row[1]) # After processing all URLs: with open('failed_urls.txt', 'w') as f: f.write('\n'.join(failed_urls))Reuse the Session (Optional): Instead of creating a new session for every
scrape()call, create one session outside the function and pass it in. This keeps connections alive and improves efficiency:# Initialize session once, before processing URLs session = requests.Session() retry = Retry(connect=5, backoff_factor=0.5, status_forcelist=[500,502,503,504]) adapter = HTTPAdapter(max_retries=retry) session.mount('http://', adapter) session.mount('https://', adapter) def scrape(page, session): try: response = session.get(page, timeout=(10,30)) response.raise_for_status() return BeautifulSoup(response.content, 'html.parser') except requests.exceptions.RequestException as e: print(f"Failed {page}: {e}") return NoneUpdate Dependencies: You're using Python 3.6 (which is end-of-life, but if you can't upgrade), make sure your
requestsandurllib3packages are up to date for the 3.6-compatible versions—older versions might have buggy retry logic.
Why This Works
- By using the session for every request, your retry rules apply to all network operations.
- Explicit timeouts prevent your script from hanging indefinitely on unresponsive servers.
- Error handling lets the script skip failed URLs instead of crashing entirely.
- Rate limiting reduces the chance of being blocked by the target server, a common cause of timeouts.
内容的提问来源于stack exchange,提问作者Small Atom

