将类字符串变量传入BeautifulSoup find_all()返回空列表问题
问题原因
你现在的核心问题是:把参数拼成字符串传给find_all(),但BeautifulSoup的find_all()需要的是独立的位置参数+关键字参数,不是拼接好的字符串。
比如手动写soup.find_all('a', href=True)时,实际传了两个参数:第一个是标签名字符串'a',第二个是关键字参数href=True;但你传self.criteria(内容为"'a', href=True")的时候,相当于只给find_all()传了一个完整字符串参数,BeautifulSoup会尝试去找标签名为"'a', href=True"的元素,网站里显然没有这种标签,所以返回空列表。
解决方案
推荐两种修改方式,按需选择:
方式1:拆分参数传递(更安全)
把criteria拆成标签名和属性字典两个独立参数,修改类的初始化和调用逻辑:
修改调用代码
import ds_gui # 拆分传标签名和属性字典 wiki_links = ds_gui.search_win( 'https://sendgrid.com', 'Window name', 'Results title', tag='a', attrs={'href': True} ) if __name__ == '__main__': wiki_links.gui.mainloop()
修改search_win类代码
import tkinter as tk from urllib.request import urlopen from bs4 import BeautifulSoup class search_win: '''class to create entire window''' # 调整初始化参数,接受tag和attrs def __init__(self, url, title, result_title, tag, attrs=None): self.url = url self.title = title self.result_title = result_title self.tag = tag self.attrs = attrs or {} # 默认空字典,兼容无属性的情况 '''creating GUI below''' self.gui = self.win_root() self.url_return = self.url_box() self.result_widget = self.results_box() self.submit_button = self.submit_button() def scrape_page(self): '''data scraping function''' #open url and scrape try: self.result_widget.delete('1.0', 'end') page = urlopen(self.url_return.get()) soup = BeautifulSoup(page, 'html.parser') # 用拆分后的参数调用find_all result = soup.find_all(self.tag, **self.attrs) except (NameError, ValueError): result = 'invalid url' print(result) finally: for res in result: # 如果想获取链接而不是文本,建议用res.get('href') final_res = res.get('href') if final_res is None: self.result_widget.insert('end', 'NO LINK\n') else: self.result_widget.insert('end', final_res + '\n') self.save_to_file() # 其余方法(win_root、url_box等)保持不变,此处省略
方式2:用eval解析字符串(快捷但有风险)
如果坚持要传字符串形式的参数,可以用eval解析字符串为合法参数,但仅适合自己可控的输入场景(避免用户输入恶意代码):
修改scrape_page方法中的find_all调用部分:
# 替换原来的result = soup.find_all(self.criteria) # 直接用eval构造调用逻辑 result = eval(f"soup.find_all({self.criteria})")
额外优化提示
你代码里用res.string获取的是标签的直接文本,很多a标签的链接地址存在href属性中,而非文本内容。如果目标是获取网站链接,建议把res.string替换为res.get('href'),这样能拿到实际的URL地址。
内容的提问来源于stack exchange,提问作者Kevin
相关产品推荐
相关产品推荐

