Python开发的Telegram机器人无法搜索IEEE Spectrum文章求助
问题排查与修复方案
核心问题分析
你的代码无法返回正确搜索结果,主要有三个关键问题:
- 搜索参数错误:IEEE Spectrum当前的搜索参数不是
keywords,而是q,原URL构造错误导致请求的页面无有效结果。 - 请求未模拟浏览器:直接用
requests.get发起请求会被网站识别为非浏览器请求,返回的页面结构可能异常。 - CSS选择器失效:网站页面结构已更新,原代码中
.search-result等选择器已不存在,无法定位文章元素。
修复后的代码
import telegram from telegram.ext import Updater, CommandHandler import requests from bs4 import BeautifulSoup def start(update, context): update.message.reply_text( "Hello! I'll help you find articles on the IEEE Spectrum website." 'Just write /search and the search keywords after that.') def search(update, context): query = " ".join(context.args) if query == "": update.message.reply_text('To search, you must enter keywords after the /search command') return url = 'https://spectrum.ieee.org' # 修正搜索参数为q,添加浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(f"{url}/search?q={query}", headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') # 修正为当前页面的文章选择器 articles = soup.select('.c-card') if len(articles) > 0: for article in articles: title_elem = article.select_one('.c-card__title a') if title_elem: title = title_elem.text.strip() href = title_elem['href'] # 处理相对链接和绝对链接 full_url = href if href.startswith('http') else f"{url}{href}" message = f'{title}\n{full_url}' update.message.reply_text(message) else: update.message.reply_text('No results were found for your request.') else: update.message.reply_text('Error when requesting IEEE Spectrum site.') bot_token = 'token' updater = Updater(token=bot_token, use_context=True) dispatcher = updater.dispatcher start_handler = CommandHandler('start', start) search_handler = CommandHandler('search', search) dispatcher.add_handler(start_handler) dispatcher.add_handler(search_handler) updater.start_polling() updater.idle()
关键修改说明
- 修正搜索URL:将
/search?keywords=改为/search?q=,匹配当前网站的搜索参数规则。 - 添加请求头:通过
User-Agent模拟浏览器请求,避免被网站拦截或返回异常页面。 - 更新CSS选择器:使用当前页面的
.c-card定位文章,.c-card__title a定位标题链接,同时增加空值判断避免报错。 - 链接处理优化:自动识别相对链接和绝对链接,确保生成的文章URL有效。
内容的提问来源于stack exchange,提问作者Steve Steve
相关产品推荐
相关产品推荐

