Python电影查询程序问题:如何仅展示网站中存在的电影?
问题描述
我尝试制作一个小型Python程序,从名为movies.txt的文本文件中读取电影名称,检查指定网站中是否存在这些电影。我编写了Python文件(main.py)和结果页面(result.html),原本应只显示网站中存在的电影,但目前却会显示文件内的所有电影名称,无论该电影是否在网站中存在。请问该如何解决这个问题?
main.py
from flask import Flask, request, render_template import requests from bs4 import BeautifulSoup app = Flask(__name__) # Open and read file with open('movies.txt', 'r') as f: movie_names = [line.strip() for line in f] @app.route('/', methods=['GET', 'POST']) def home(): if request.method == 'POST': # Get file file = request.files['file'] # Read file movie_names = [] for line in file: movie_names.append(line.strip()) # Websites for movies movie_info = {} url = ['https://www.websiteone.com/', 'https://www.websitetwo.com', 'https://www.websitethree.com/', 'https://www.websitefour.com', 'https://www.websitefive.com', 'https://www.websitesix.com'] for name in movie_names: for url in urls: response = requests.get(url + name) soup = BeautifulSoup(response.content, 'html.parser') info = soup.find('div', {'class': 'movie-info'}) movie_info[name] = info.text if info else 'Not found' return render_template('result.html', movie_info=movie_info) return render_template('home.html') if __name__ == '__main__': app.run(debug=True)
result.html
<!DOCTYPE html> <html> <head> <title>Movie Information</title> </head> <body> <h1>Movie Information</h1> {% if movie_info %} <table> <thead> <tr> <th>Movie Name</th> <th>Website</th> </tr> </thead> <tbody> {% for name, website in movie_info.items() %} <tr> <td>{{ name }}</td> <td>{{ website }}</td> </tr> {% endfor %} </tbody> </table> {% else %} <p>No movie found.</p> {% endif %} </body> </html>
解决方案
你的代码存在几个关键问题,导致所有电影都被显示:
- 变量名错误
定义网站列表时用了url,但循环遍历网站时用的是urls,这会直接触发NameError。需要把列表名统一改成urls:
urls = ['https://www.websiteone.com/', 'https://www.websitetwo.com', 'https://www.websitethree.com/', 'https://www.websitefour.com/', 'https://www.websitefive.com', 'https://www.websitesix.com']
- 逻辑错误:覆盖结果且未过滤不存在的电影
- 当前代码每遍历一个网站,就会覆盖
movie_info[name]的取值,最后只保留最后一个网站的检查结果; - 无论电影是否找到,都会强制加入
movie_info字典,导致所有电影都会被渲染到页面上。
修改核心逻辑,只保留找到的电影:
movie_info = {} urls = ['https://www.websiteone.com/', 'https://www.websitetwo.com', 'https://www.websitethree.com/', 'https://www.websitefour.com/', 'https://www.websitefive.com', 'https://www.websitesix.com'] for name in movie_names: found = False for site_url in urls: try: # 加入异常处理,避免网站访问失败导致程序崩溃 response = requests.get(site_url + name) response.raise_for_status() soup = BeautifulSoup(response.content, 'html.parser') info = soup.find('div', {'class': 'movie-info'}) if info: # 记录找到电影的网站(或电影信息,根据需求调整) movie_info[name] = site_url found = True break # 找到后停止遍历其他网站,提升效率 except requests.exceptions.RequestException as e: print(f"访问{site_url}出错: {e}") # 未找到的电影不加入字典,自然不会被显示 return render_template('result.html', movie_info=movie_info)
- 模板适配
修改后movie_info的value存储的是电影存在的网站地址,和你现有模板的<th>Website</th>对应。如果需要显示电影详情,可将movie_info[name]改为info.text,逻辑保持一致即可。
内容的提问来源于stack exchange,提问作者Salieri Seppuku
相关产品推荐
相关产品推荐

