使用Python+Selenium爬取Marathi Matrimony网站时遇Out of Memory Error求助
爬取Marathi Matrimony网站时的内存溢出问题及解决思路
问题背景
- 目标爬取网站共365页内容,但仅能成功爬取前130页,之后网站抛出Out of Memory Error并停止响应
- 刷新页面会重置回第1页,需重新点击130次「下一页」按钮才能继续爬取,效率极低
现有爬取代码
# Create ChromeOptions and set incognito mode chrome_options = Options() chrome_options.add_argument('--incognito') driver = webdriver.Chrome(options=chrome_options) driver.get('https://matches.marathimatrimony.com/') username = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[name='MIDP']"))) password = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[name='PASSWORD2']"))) #enter username and password username.clear() username.send_keys("EMAIL") password.clear() password.send_keys("PASSWORD") # Wait for the login button to be visible button = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "input[type='submit'][value='LOGIN']"))) # Click the login button button.click() sleep(2) # Find the element using the CSS selector ad_close_button = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "span > img[alt='Close button']"))) # Click the element ad_close_button.click() sleep(2) # Find the element using the CSS selector all_matches = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "a[routerlink='/listing']"))) # Click the element all_matches.click() # Create a DataFrame to store the data columns = ["Profile Name", "Profile Href", "Profile ID", "Page Number"] data2 = [] current_page_number = 1 last_page_scraped = 0 count = 1 while current_page_number <= 375: if current_page_number <= last_page_scraped: print("Current page is not equal to the last page scraped") while current_page_number <= last_page_scraped: next_page_element = driver.find_element(By.XPATH, "//li[@class='pagination-next']//a") driver.execute_script("arguments[0].click();", next_page_element) current_page_element = driver.find_element(By.XPATH, "//li[contains(@class, 'current')]/span[2]") current_page_text = current_page_element.text current_page_number = re.search(r'\d+', current_page_text).group() current_page_number = int(current_page_number) print(f"Moved to page: {current_page_number}") sleep(2) else: print(f"Scraping page number: {current_page_number}") # Find the profile elements profile_elements = driver.find_elements(By.CSS_SELECTOR, "div.listingMatchCard") # Loop through each profile element for profile in profile_elements: try: profile_name_element = profile.find_element(By.CSS_SELECTOR, "a.clr-black1.text-decoration-none.col-md-12.pl-0") profile_name = profile_name_element.text profile_href = profile_name_element.get_attribute("href") profile_id_element = profile.find_element(By.CSS_SELECTOR, "a.cursor-pointer.outline-none.text-decoration-none.clr-grey2") profile_id = profile_id_element.get_attribute("href").split("/")[-1] data2.append([profile_name, profile_href, profile_id, current_page_number]) count += 1 except NoSuchElementException: print("Error occurred while scraping a profile. Skipping...") last_page_scraped = current_page_number print(f"Successfully scraped page number: {last_page_scraped}") if current_page_number % 5 == 0: # Save the scraped data every 5 pages df2 = pd.DataFrame(data2, columns=columns) filename = f"scraped_profiles/scraped_profiles_page_{current_page_number}.csv" df2.to_csv(filename, index=False) try: next_page_element = WebDriverWait(driver, 20).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "li.pagination-next a"))) driver.execute_script("arguments[0].click();", next_page_element) except TimeoutException: break sleep(2) current_page_element = driver.find_element(By.XPATH, "//li[contains(@class, 'current')]/span[2]") current_page_text = current_page_element.text current_page_number = re.search(r'\d+', current_page_text).group() current_page_number = int(current_page_number) print(f"The current page is: {current_page_number}") # Save the final scraped data df2 = pd.DataFrame(data2, columns=columns) filename = f"final_scraped_profiles.csv" df2.to_csv(filename, index=False)
尝试过的分页跳转方法
曾尝试直接跳转到目标页码,但仍会在点击约130次「下一页」后出现内存问题,代码如下:
target_page_number = 135 while True: try: current_page_element = driver.find_element(By.XPATH, "//li[contains(@class, 'current')]/span[2]") current_page_text = current_page_element.text current_page_number = re.search(r'\d+', current_page_text).group() current_page_number = int(current_page_number) if current_page_number >= target_page_number: break except NoSuchElementException: break print(f'current page is:{current_page_number}',end='\r') next_page_element = driver.find_element(By.XPATH, "//li[@class='pagination-next']//a") driver.execute_script("arguments[0].click();", next_page_element) sleep(2)
解决思路与优化建议
1. 定期重启浏览器释放内存
Selenium长时间运行会积累大量内存占用,建议每爬取固定页数(如50页)就关闭当前浏览器实例,重新登录并跳转到对应页码继续:
- 每次重启前将当前爬取到的页码写入本地文件(如
last_page.txt),下次启动时读取该文件直接跳转 - 核心逻辑示例:
import os # 读取上次爬取的页码 if os.path.exists('last_page.txt'): with open('last_page.txt', 'r') as f: last_page_scraped = int(f.read().strip()) else: last_page_scraped = 0 # 每爬取50页重启一次 if current_page_number % 50 == 0: # 保存当前页码 with open('last_page.txt', 'w') as f: f.write(str(current_page_number)) # 关闭浏览器 driver.quit() # 重新初始化driver、登录、跳转到current_page_number # 复用已有的登录和分页跳转逻辑
2. 优化内存中的数据存储
- 避免全局列表
data2存储过多数据,每爬取20-30页就将数据写入CSV并清空列表:if current_page_number % 20 == 0: df2 = pd.DataFrame(data2, columns=columns) df2.to_csv(f"scraped_profiles/scraped_profiles_page_{current_page_number}.csv", index=False) data2.clear() # 清空列表释放内存 import gc gc.collect() # 强制触发垃圾回收 - 爬取完每页后,主动清空元素引用:
# 爬取完一页后执行 profile_elements = None gc.collect()
3. 启用无头浏览器模式
Chrome的UI渲染会占用大量内存,启用无头模式可显著降低内存消耗:
chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') # 解决Linux环境下的内存限制问题
4. 绕过前端渲染,直接调用API接口
- 打开浏览器开发者工具(F12),切换到「网络」标签,点击「下一页」时观察XHR/fetch请求,找到返回分页数据的API接口
- 使用
requests库直接请求该接口,无需加载整个页面,内存占用会大幅降低 - 注意复制登录后的Cookie、Authorization等身份验证参数到请求头中,确保接口请求有效
5. 优化分页跳转逻辑
- 检查页面是否存在可直接输入页码的输入框,若有则直接输入目标页码并提交,替代多次点击「下一页」:
from selenium.webdriver.common.keys import Keys # 假设存在页码输入框,需根据实际页面调整选择器 page_input = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input.pagination-page-input"))) page_input.clear() page_input.send_keys(target_page_number) page_input.send_keys(Keys.ENTER) - 检查URL是否支持分页参数(如
?page=130),若支持则直接拼接URL跳转:driver.get(f"https://matches.marathimatrimony.com/listing?page={target_page_number}")
内容的提问来源于stack exchange,提问作者Prathamesh Sawant
相关产品推荐
相关产品推荐

