使用Python Selenium爬取维基百科时遇AttributeError问题求助
问题解决:Selenium AttributeError 修复
错误原因
你遇到的AttributeError是因为Selenium 4.0及以上版本已经废弃了find_element_by_xpath、find_element_by_css_selector这类旧的元素定位方法,统一使用find_element()和find_elements()方法,通过by和value参数指定定位策略和表达式。
修复步骤
1. 导入必要类
需要从selenium.webdriver.common.by导入By类用于指定定位策略;同时推荐导入WebDriverWait和expected_conditions做显式等待,避免页面未加载完成导致的元素查找失败。
2. 替换旧定位代码
把原来的page_container = driver.find_element_by_xpath('//*[@class="mw-page-container"]')替换为新语法。
3. 添加显式等待(可选但强烈推荐)
维基百科页面加载需要时间,直接查找元素容易失败,添加显式等待可确保元素渲染完成后再操作。
修复后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize the webdriver driver = webdriver.Edge() WIKI = "https://en.wikipedia.org/wiki/India" driver.get(WIKI) # 显式等待页面容器元素加载完成,最多等待10秒 page_container = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//*[@class="mw-page-container"]')) ) # Find all elements of interest inside the page container elements = page_container.find_elements(by=By.CSS_SELECTOR, value='h1, h2, h3, h4, h5, h6, p, ul, ol, li') # Save the content to a text file with open('site_content.txt', 'w', encoding='utf-8') as file: for element in elements: tag_name = element.tag_name if tag_name.startswith('h'): # Check if it's a heading file.write(f"-----### {element.text} ###-----\n") elif tag_name == 'p': # Check if it's a paragraph file.write(element.text + '\n') elif tag_name in ['ul', 'ol']: # Check if it's an unordered/ordered list list_items = element.find_elements(by=By.TAG_NAME, value='li') for item in list_items: file.write(f"- {item.text}\n") # Don't forget to close the webdriver driver.quit() print("Content saved to 'site_content.txt'")
额外说明
- 所有旧的
find_element_by_*方法都需要替换为find_element(by=By.XXX, value='表达式'),比如find_element_by_id('xxx')要改成find_element(by=By.ID, value='xxx')。 - 显式等待的超时时间可根据网络情况调整,示例中设置为10秒。
内容的提问来源于stack exchange,提问作者Keshav Verma
相关产品推荐
相关产品推荐

