Python烂番茄Top100电影爬虫报错:AttributeError无getText属性
Ah, I spot the issue immediately! That AttributeError pops up because you accidentally used the homepage's soup object instead of the movie detail page's soup2 object when trying to fetch the audience rating span. Since the homepage doesn't contain that specific element, soup.find(...) returns None, and calling .getText() on a None value throws the error.
Here's the corrected code with that critical fix, plus a small robustness improvement for URL construction:
from urllib.request import urlopen from bs4 import BeautifulSoup import random url = "https://www.rottentomatoes.com/top/bestofrt/" page = urlopen(url) soup = BeautifulSoup(page, 'html.parser') table = soup.find("table", attrs={"class":"table"}) links = table.findAll("a") titles = [] hyperLinks = [] for link in links: titles.append(link.getText().strip()) hyperLinks.append(link["href"]) choice = random.randint(0, len(hyperLinks) - 1) # More robust way to build the detail page URL detail_url = f"{url.rstrip('/top/bestofrt/')}{hyperLinks[choice]}" page2 = urlopen(detail_url) soup2 = BeautifulSoup(page2, "html.parser") # Use soup2 (detail page's parsed content) instead of soup! span = soup2.find("span", attrs={"class":"mop-ratings-wrap__percentage"}) print(titles[choice]) print(span.getText().strip())
Key Fixes & Notes:
- Core Error Fix: Swapped
soup.find(...)withsoup2.find(...)—soup2holds the parsed HTML of the specific movie's detail page, which actually contains the rating span you're targeting. - URL Robustness: Replaced the hardcoded slice
url[:-14]withurl.rstrip('/top/bestofrt/')to avoid breaking the URL if the base path ever changes.
A quick heads-up: Rotten Tomatoes occasionally updates their page structure, so if this code stops working later, you'll want to re-inspect the detail page's HTML to confirm the rating element's class name is still valid.
内容的提问来源于stack exchange,提问作者Night Balasingham

