Python Requests库iter_content下载文件异常排查与解决
Hey there! Let me share this fun little project I built for my girlfriend, along with the bugs I fixed and what I learned about that tricky NoneType is not iterable error.
Project Background
I'm building a program that automatically downloads cute images of puppies, kittens, and other baby animals (with plans to expand it later to download PDFs, books, movies, and more). The core idea is to pull random images from a list of curated websites, but I ran into two big issues:
- Some images would download incomplete (tiny file sizes that couldn't be opened)
- Occasional
NoneType is not iterableerrors crashing the script
Original Code
Here's the initial version of the code that had the issues:
#! python3 import os, requests, bs4, random, shelve, wget def create_folder(): #Creates the folder 'Puppies' print('Comprobando pelusitas y pulguitas...') os.makedirs('Puppies', exist_ok=True) def download_image(url, request_response): #Saves the image in the folder 'Puppies' image_file = open(os.path.join('Puppies', os.path.basename(url)), 'wb') for chunk in request_response.iter_content(chunk_size=1024): image_file.write(chunk) image_file.flush() image_file.close() def create_saves(saves_list): #Creates a data file, if not yet created, storing the pictures already downloaded. if os.path.isfile('C:\\Users\\usuario\\Documents\\santi\\saves.dat') == False: page_files = shelve.open('saves') page_files['saves'] = saves_list page_files.close() def look_images(pages_list, saves_list): for i in pages_list: #For every item in the pages_list (for every page...) res = requests.get(i) res.raise_for_status() soup = bs4.BeautifulSoup(res.text, "html.parser") object = soup.select('img') object_number = len(object) random_number = random.randint(0, int(object_number)) #Creates a random number between 0 and the number of images the object variable found. print('\nSe encontraron {} cachorritos potenciales...'.format(int(object_number))) try: image_url = object[random_number].get('srcset') except IndexError: print('No se hallaron vauitos, se continuará la búsqueda...') print('\nEvaluando colmillitos y patitas...') page_files = shelve.open('saves') #Opens the saves data file. while str(image_url) in open('saves.dat', encoding='Latin-1').read(): #While the image found is stored in the saves.dat file, select another image. try: image_url = object[random.randint(0, int(object_number))].get('src') except IndexError: continue while not '.jpg' in str(image_url): try: image_url = object[random.randint(0, int(object_number))].get('src') if str(image_url) in open('saves.dat', encoding='Latin-1').read(): image_url = object[random.randint(0, int(object_number))].get('src') continue except IndexError: print('No se hallaron vauitos, se continuará la búsqueda...') continue print('\nSe encontraron vauitos...') print('\nAdoptando cachorrito...') if str(image_url).endswith('.jpg 2x'): #Lot of images were downloaded as '.jpg 2x', so I made this if statement to erase the ' 2x' final part. image_url = str(image_url.replace(' ', '')[:-2]) response = requests.get(image_url, stream=True) res.raise_for_status() saves_list.append(image_url) #Adds the image to the save_list, which is then saved on the .dat 'saves' file. download_image(image_url, response) print('¡Cachorrito adoptado!') page_files[image_url] = page_files #Saves the image url on the 'saves'.dat file page_files.close() def get_page(): page = ['https://pixabay.com/es/photos/puppy/', 'https://www.petsworld.in/blog/cute-pictures-of-puppies-and-kittens-together.html', 'https://pixabay.com/es/photos/bear%20cubs/', 'https://pixabay.com/es/photos/?q=cute+little+animals&hp=&image_type=photo&order=popular&cat=', 'https://pixabay.com/es/photos/?q=baby+cows&hp=&image_type=photo&order=popular&cat=', 'https://www.boredpanda.com/cute-baby-animals/', 'http://abduzeedo.com/node/74367'] alr_dow = [] create_folder() create_saves(alr_dow) look_images(page, alr_dow) get_page()
Troubleshooting & Key Fixes
After digging into the code, here's what I found and fixed:
- Misused Error Handling: The biggest culprit for incomplete downloads was a typo in
look_images: I calledres.raise_for_status()after fetching the image URL, butreswas the response from loading the webpage—not the image itself. Switching this toresponse.raise_for_status()let me catch errors like blocked ports or firewall restrictions that were causing partial image downloads. - Unreliable Download Method: Replaced the manual chunk-writing
download_imagefunction withshutil.copyfileobj—it's more robust for streaming files and handles edge cases better. - Encapsulated URL Logic: Pulled the random valid URL selection into its own
get_random_urlfunction to clean up the code and make it easier to debug. - Relative URL Handling: Added a check to prepend
https://to image URLs that didn't start with it, fixing failed requests to images with relative paths. - NoneType Prevention: Converted variables to strings where needed to avoid the
NoneType is not iterableerror (more on that below).
Optimized Code
Here's the fixed version of the script:
#! python3 import os, requests, bs4, random, shelve, shutil def create_folder(): print('Comprobando pelusitas y pulguitas...') os.makedirs('Puppies', exist_ok=True) def download_image_shutil(url, request_response): local_filename = url.split('/')[-1] with open(os.path.join('Puppies', local_filename), 'wb') as f: shutil.copyfileobj(request_response.raw, f, [False]) return local_filename def create_saves(saves_list): if os.path.isfile('C:\\Users\\usuario\\Documents\\santi\\saves.dat') == False: page_files = shelve.open('saves') page_files['saves'] = saves_list page_files.close() def get_random_url(web_object, web_object_number): while True: try: random_number = random.randint(0, web_object_number) url_variable = web_object[random_number].get('src') if not '.jpg' in str(url_variable): continue elif str(url_variable) in open('saves.dat', encoding='Latin-1').read(): continue print('Finish') print(url_variable) return url_variable break except IndexError: print('Algo salió mal: reanundando búsqueda...') continue def look_images(pages_list, saves_list): for i in pages_list: res = requests.get(i) res.raise_for_status() soup = bs4.BeautifulSoup(res.text, "html.parser") object = soup.select('img') object_number = len(object) print('\nSe encontraron {} cachorritos potenciales...'.format(int(object_number))) print('\nEvaluando colmillitos y patitas...') page_files = shelve.open('saves') image_url = get_random_url(object, object_number) print('\nSe encontraron vauitos...') print('\nAdoptando cachorrito...') if str(image_url).endswith('.jpg 2x'): image_url = str(image_url.replace(' ', '')[:-2]) if not str(image_url).startswith('https://'): image_url = 'https://' + str(image_url) response = requests.get(image_url, stream=True) response.raise_for_status() saves_list.append(image_url) download_image_shutil(image_url, response) print('¡Cachorrito adoptado!') page_files[image_url] = page_files page_files.close() def get_page(): page = ['https://pixabay.com/es/photos/puppy/', 'https://www.petsworld.in/blog/cute-pictures-of-puppies-and-kittens-together.html', 'https://pixabay.com/es/photos/bear%20cubs/', 'https://pixabay.com/es/photos/?q=cute+little+animals&hp=&image_type=photo&order=popular&cat=', 'https://pixabay.com/es/photos/?q=baby+cows&hp=&image_type=photo&order=popular&cat=', 'https://www.boredpanda.com/cute-baby-animals/'] alr_dow = [] create_folder() create_saves(alr_dow) look_images(page, alr_dow) get_page()
Why the "NoneType is not iterable" Error Happens
Let's break this down simply:
When you use methods like element.get('src') on a BeautifulSoup tag, if that tag doesn't have the src attribute, it returns None instead of a string.
If you then try to do something like '.jpg' in url_variable where url_variable is None, Python throws the NoneType is not iterable error because it can't check if a substring exists in a non-string (non-iterable) value like None.
By converting url_variable to a string with str(url_variable), even if it's None, it becomes the string "None"—which is iterable, so the in check works without crashing the script. That small conversion prevents the error and lets the loop keep looking for a valid image URL.
内容的提问来源于stack exchange,提问作者lafinur

