Python电影爬虫问题:同时打印两类信息时无法获取全部数据
Hey there! Let's break down why your scraper works when printing movie or addinfo alone, but fails to show all data when you print them together. This is a super common pitfall in web scraping—here are the most likely causes and fixes:
1. You're Overwriting Variables Instead of Storing All Entries
If you’re using a single variable (like movie or addinfo) in a loop and just updating it each time, you’ll only end up with the last item’s data when you print later. For example:
# ❌ Bad: Variables get overwritten each loop movie = "" addinfo = "" for item in movie_elements: movie = item.find("h3").text link = item.find("a")["href"] addinfo = requests.get(link).text # Only saves the last link's info print(movie, addinfo) # Only shows the final movie and its info
Fix: Store each pair of data in a list of dictionaries so you keep every entry:
# ✅ Good: Save all entries to a list scraped_data = [] for item in movie_elements: # Extract movie name movie_name = item.find("h3").text.strip() # Fetch and parse additional info from the link info_link = item.find("a")["href"] try: response = requests.get(info_link) response.raise_for_status() # Catch HTTP errors additional_info = response.text.strip() # Or parse with BeautifulSoup here except Exception as e: additional_info = f"Error fetching info: {str(e)}" # Add the pair to your list scraped_data.append({ "movie": movie_name, "addinfo": additional_info }) # Now print all entries for entry in scraped_data: print(f"Movie: {entry['movie']}") print(f"Additional Info: {entry['addinfo']}\n")
2. You're Exhausting an Iterator
If movie or addinfo comes from an iterator (like a generator or a find_all result treated as an iterator instead of a list), the first time you loop through it to print, you’ll deplete it. The second loop will have nothing left.
For example, if you do:
# ❌ Bad: Iterators get exhausted after first use movies = soup.find_all("h3") # This is a ResultSet, but if converted to a generator... addinfo_links = soup.find_all("a") # First loop uses up the movies iterator for movie in movies: print(movie.text) # Second loop has no movies left! for movie, link in zip(movies, addinfo_links): print(movie.text, requests.get(link["href"]).text)
Fix: Convert iterators to lists upfront so you can reuse them:
# ✅ Good: Convert to lists to preserve all items movies = list(soup.find_all("h3")) addinfo_links = list(soup.find_all("a")) # Now you can loop through them multiple times without issues for movie, link in zip(movies, addinfo_links): print(f"Movie: {movie.text.strip()}") print(f"Info: {requests.get(link['href']).text.strip()}\n")
3. Asynchronous/Request Timing Issues (Less Likely but Possible)
If you’re using async requests (like with aiohttp) and not properly waiting for each request to finish before moving to the next movie, you might end up with mismatched data (e.g., the last request’s info gets paired with all movies).
Fix: Ensure you’re awaiting each request inside the loop, or use a synchronous approach like requests if you don’t need async speed.
Quick Check List
- Are you storing every movie-info pair, not just overwriting variables?
- Are you using lists instead of one-time iterators for your elements?
- If using async, are you properly syncing requests with movie entries?
By following these steps, you should be able to print all your movie names and their corresponding additional info without missing any data!
内容的提问来源于stack exchange,提问作者SIM

