构建方言抓取网页爬虫时遭遇AttributeError: module 'bs4.element' has no attribute 'find'错误的求助
解决AttributeError: module 'bs4.element' has no attribute 'find'的问题
嘿,这个错误其实是个很容易犯的小疏忽——你把模块和DOM元素对象搞混啦!
错误原因分析
你导入了bs4.element模块,然后在循环里写了element.find('h4'),但element是整个模块,它根本没有find方法。真正拥有find方法的是你循环里的每个words变量——也就是遍历soup.find('div', class_='panel-body')得到的单个DOM元素节点。
修正后的代码
这里给你调整好的代码,还加了一些容错处理,避免遇到非标签节点时报错:
from bs4 import BeautifulSoup, element import pandas as pd import requests capture = requests.get('http://wiwords.com/dictionary/').text results = [] soup = BeautifulSoup(capture, 'lxml') # 先定位到目标面板 panel_body = soup.find('div', class_='panel-body') # 遍历面板下的所有标签节点,过滤掉文本/换行这类非标签内容 for word_element in panel_body.find_all(): # 用当前遍历到的元素调用find方法,而不是element模块 extracted_h4 = word_element.find('h4') if extracted_h4: # 先确认找到h4元素 clean_text = extracted_h4.text.strip() # 确保文本非空且没有重复添加 if clean_text and clean_text not in results: results.append(clean_text) # 查看结果 print(results)
更高效的写法(可选)
如果你的目标就是抓取页面上所有的h4标签内容,其实可以直接跳过面板遍历,一步到位:
from bs4 import BeautifulSoup import requests capture = requests.get('http://wiwords.com/dictionary/').text soup = BeautifulSoup(capture, 'lxml') # 直接抓取所有h4标签,提取文本并去重 h4_texts = [tag.text.strip() for tag in soup.find_all('h4') if tag.text.strip()] results = list(set(h4_texts)) # 转集合去重,再转回列表 print(results)
内容的提问来源于stack exchange,提问作者Keston Smith
相关产品推荐
相关产品推荐

