Selenium中使用getAttribute()无结果?爬取黄页邮箱报错排查
解决爬取澳大利亚黄页时获取邮箱的TypeError问题
嘿,我来帮你搞定这个爬取黄页时遇到的TypeError: 'NoneType' object is not callable问题!这个错误其实是两个小问题叠加导致的,咱们一步步拆解解决:
错误原因分析
- 工具方法混淆:你把Selenium WebElement的方法和BeautifulSoup的Tag对象搞混啦!BeautifulSoup解析出来的元素根本没有
getAttribute()这个方法——它获取属性的正确姿势是用字典索引或者get()方法。 - 空值未处理:不是每一家披萨餐厅都会公开邮箱,当
item.find('a', class_='contact contact-main contact-email ')找不到对应元素时,会返回None,这时候你直接调用方法自然就报错了。
修正后的代码
我把你的代码调整了一下,加上了空值判断和正确的属性获取方式,还补充了一些细节优化:
import csv from bs4 import BeautifulSoup import requests from selenium import webdriver from selenium.webdriver.common.by import By url = "https://www.yellowpages.com.au/search/listings?clue=Pizza+Restaurants&locationClue=Sydney+CBD%2C+NSW&lat=&lon=" # 如果你用的是Selenium 4.x版本,建议改用Service类(避免executable_path弃用警告) # from selenium.webdriver.chrome.service import Service # service = Service("/usr/local/share/chromedriver") # driver = webdriver.Chrome(service=service) driver = webdriver.Chrome(executable_path="/usr/local/share/chromedriver") driver.get(url) pageSource = driver.page_source bsObj = BeautifulSoup(pageSource, 'lxml') # 先确认外层容器存在,避免后续findAll直接出错 outer_container = bsObj.find('div', {'class': 'flow-layout outside-gap-large inside-gap inside-gap-large vertical'}) if outer_container: items = outer_container.findAll('div', class_='cell in-area-cell find-show-more-trial middle-cell') for item in items: # 查找邮箱链接,先判断是否存在 email_link = item.find('a', class_='contact contact-main contact-email') if email_link: # 用BeautifulSoup的get方法获取属性,安全又省心 email = email_link.get('data-email') print(email) else: print("这家餐厅没有公开邮箱哦") else: print("哎呀,没找到餐厅列表容器,可能页面结构更新啦,得重新检查一下class名") # 记得关闭浏览器,避免资源浪费 driver.quit()
额外注意事项
- 属性获取小技巧:用
tag.get('属性名')比直接tag['属性名']更安全,就算属性不存在也只会返回None,不会抛出KeyError。 - 页面结构变动:黄页网站的class名可能会不定期更新,如果之后还是找不到元素,记得打开浏览器开发者工具,重新核对当前页面的元素结构。
- Selenium版本兼容:如果你用的是Selenium 4.x及以上版本,
executable_path参数已经被弃用了,建议改用代码里注释的Service类写法,避免警告。
内容的提问来源于stack exchange,提问作者Mukul Agrawal
相关产品推荐
相关产品推荐

