You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python电影爬虫问题:同时打印两类信息时无法获取全部数据

Troubleshooting Your Python Movie Scraper: Missing Info When Printing Both Movie Names & Details

Hey there! Let's break down why your scraper works when printing movie or addinfo alone, but fails to show all data when you print them together. This is a super common pitfall in web scraping—here are the most likely causes and fixes:

1. You're Overwriting Variables Instead of Storing All Entries

If you’re using a single variable (like movie or addinfo) in a loop and just updating it each time, you’ll only end up with the last item’s data when you print later. For example:

# ❌ Bad: Variables get overwritten each loop
movie = ""
addinfo = ""
for item in movie_elements:
    movie = item.find("h3").text
    link = item.find("a")["href"]
    addinfo = requests.get(link).text  # Only saves the last link's info

print(movie, addinfo)  # Only shows the final movie and its info

Fix: Store each pair of data in a list of dictionaries so you keep every entry:

# ✅ Good: Save all entries to a list
scraped_data = []
for item in movie_elements:
    # Extract movie name
    movie_name = item.find("h3").text.strip()
    
    # Fetch and parse additional info from the link
    info_link = item.find("a")["href"]
    try:
        response = requests.get(info_link)
        response.raise_for_status()  # Catch HTTP errors
        additional_info = response.text.strip()  # Or parse with BeautifulSoup here
    except Exception as e:
        additional_info = f"Error fetching info: {str(e)}"
    
    # Add the pair to your list
    scraped_data.append({
        "movie": movie_name,
        "addinfo": additional_info
    })

# Now print all entries
for entry in scraped_data:
    print(f"Movie: {entry['movie']}")
    print(f"Additional Info: {entry['addinfo']}\n")

2. You're Exhausting an Iterator

If movie or addinfo comes from an iterator (like a generator or a find_all result treated as an iterator instead of a list), the first time you loop through it to print, you’ll deplete it. The second loop will have nothing left.

For example, if you do:

# ❌ Bad: Iterators get exhausted after first use
movies = soup.find_all("h3")  # This is a ResultSet, but if converted to a generator...
addinfo_links = soup.find_all("a")

# First loop uses up the movies iterator
for movie in movies:
    print(movie.text)

# Second loop has no movies left!
for movie, link in zip(movies, addinfo_links):
    print(movie.text, requests.get(link["href"]).text)

Fix: Convert iterators to lists upfront so you can reuse them:

# ✅ Good: Convert to lists to preserve all items
movies = list(soup.find_all("h3"))
addinfo_links = list(soup.find_all("a"))

# Now you can loop through them multiple times without issues
for movie, link in zip(movies, addinfo_links):
    print(f"Movie: {movie.text.strip()}")
    print(f"Info: {requests.get(link['href']).text.strip()}\n")

3. Asynchronous/Request Timing Issues (Less Likely but Possible)

If you’re using async requests (like with aiohttp) and not properly waiting for each request to finish before moving to the next movie, you might end up with mismatched data (e.g., the last request’s info gets paired with all movies).

Fix: Ensure you’re awaiting each request inside the loop, or use a synchronous approach like requests if you don’t need async speed.

Quick Check List

  • Are you storing every movie-info pair, not just overwriting variables?
  • Are you using lists instead of one-time iterators for your elements?
  • If using async, are you properly syncing requests with movie entries?

By following these steps, you should be able to print all your movie names and their corresponding additional info without missing any data!

内容的提问来源于stack exchange,提问作者SIM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:29:54