You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium爬取Marathi Matrimony网站时遇Out of Memory Error求助

爬取Marathi Matrimony网站时的内存溢出问题及解决思路

问题背景

  • 目标爬取网站共365页内容,但仅能成功爬取前130页,之后网站抛出Out of Memory Error并停止响应
  • 刷新页面会重置回第1页,需重新点击130次「下一页」按钮才能继续爬取,效率极低

现有爬取代码

# Create ChromeOptions and set incognito mode
chrome_options = Options()
chrome_options.add_argument('--incognito')

driver = webdriver.Chrome(options=chrome_options)
driver.get('https://matches.marathimatrimony.com/')

username = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[name='MIDP']")))
password = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[name='PASSWORD2']")))
                                           
#enter username and password
username.clear()
username.send_keys("EMAIL")
password.clear()
password.send_keys("PASSWORD")

# Wait for the login button to be visible
button = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "input[type='submit'][value='LOGIN']")))

# Click the login button
button.click()

sleep(2)

# Find the element using the CSS selector
ad_close_button = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "span > img[alt='Close button']")))

# Click the element
ad_close_button.click()

sleep(2)

# Find the element using the CSS selector
all_matches = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "a[routerlink='/listing']")))

# Click the element
all_matches.click()

# Create a DataFrame to store the data
columns = ["Profile Name", "Profile Href", "Profile ID", "Page Number"]
data2 = []
current_page_number = 1
last_page_scraped = 0
count = 1

while current_page_number <= 375:
    if current_page_number <= last_page_scraped:
        print("Current page is not equal to the last page scraped")
        while current_page_number <= last_page_scraped:
            next_page_element = driver.find_element(By.XPATH, "//li[@class='pagination-next']//a")
            driver.execute_script("arguments[0].click();", next_page_element)
            current_page_element = driver.find_element(By.XPATH, "//li[contains(@class, 'current')]/span[2]")
            current_page_text = current_page_element.text
            current_page_number = re.search(r'\d+', current_page_text).group()
            current_page_number = int(current_page_number)
            print(f"Moved to page: {current_page_number}")
        sleep(2)
    else:
        print(f"Scraping page number: {current_page_number}")
        # Find the profile elements
        profile_elements = driver.find_elements(By.CSS_SELECTOR, "div.listingMatchCard")
        # Loop through each profile element
        for profile in profile_elements:
            try:
                profile_name_element = profile.find_element(By.CSS_SELECTOR, "a.clr-black1.text-decoration-none.col-md-12.pl-0")
                profile_name = profile_name_element.text
                profile_href = profile_name_element.get_attribute("href")
                profile_id_element = profile.find_element(By.CSS_SELECTOR, "a.cursor-pointer.outline-none.text-decoration-none.clr-grey2")
                profile_id = profile_id_element.get_attribute("href").split("/")[-1]
                data2.append([profile_name, profile_href, profile_id, current_page_number])
                count += 1
            except NoSuchElementException:
                print("Error occurred while scraping a profile. Skipping...")
        last_page_scraped = current_page_number
        print(f"Successfully scraped page number: {last_page_scraped}")
        if current_page_number % 5 == 0:
            # Save the scraped data every 5 pages
            df2 = pd.DataFrame(data2, columns=columns)
            filename = f"scraped_profiles/scraped_profiles_page_{current_page_number}.csv"
            df2.to_csv(filename, index=False)
        try:
            next_page_element = WebDriverWait(driver, 20).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "li.pagination-next a")))
            driver.execute_script("arguments[0].click();", next_page_element)
        except TimeoutException:
            break
        sleep(2)
        current_page_element = driver.find_element(By.XPATH, "//li[contains(@class, 'current')]/span[2]")
        current_page_text = current_page_element.text
        current_page_number = re.search(r'\d+', current_page_text).group()
        current_page_number = int(current_page_number)
        print(f"The current page is: {current_page_number}")

# Save the final scraped data
df2 = pd.DataFrame(data2, columns=columns)
filename = f"final_scraped_profiles.csv"
df2.to_csv(filename, index=False)

尝试过的分页跳转方法

曾尝试直接跳转到目标页码,但仍会在点击约130次「下一页」后出现内存问题,代码如下:

target_page_number = 135

while True:
    try:
        current_page_element = driver.find_element(By.XPATH, "//li[contains(@class, 'current')]/span[2]")
        current_page_text = current_page_element.text
        current_page_number = re.search(r'\d+', current_page_text).group()
        current_page_number = int(current_page_number)
        if current_page_number >= target_page_number:
            break
    except NoSuchElementException:
        break
    
    print(f'current page is:{current_page_number}',end='\r')
    next_page_element = driver.find_element(By.XPATH, "//li[@class='pagination-next']//a")
    driver.execute_script("arguments[0].click();", next_page_element)
    sleep(2)

解决思路与优化建议

1. 定期重启浏览器释放内存

Selenium长时间运行会积累大量内存占用,建议每爬取固定页数(如50页)就关闭当前浏览器实例,重新登录并跳转到对应页码继续:

  • 每次重启前将当前爬取到的页码写入本地文件(如last_page.txt),下次启动时读取该文件直接跳转
  • 核心逻辑示例:
    import os
    # 读取上次爬取的页码
    if os.path.exists('last_page.txt'):
        with open('last_page.txt', 'r') as f:
            last_page_scraped = int(f.read().strip())
    else:
        last_page_scraped = 0
    
    # 每爬取50页重启一次
    if current_page_number % 50 == 0:
        # 保存当前页码
        with open('last_page.txt', 'w') as f:
            f.write(str(current_page_number))
        # 关闭浏览器
        driver.quit()
        # 重新初始化driver、登录、跳转到current_page_number
        # 复用已有的登录和分页跳转逻辑
    

2. 优化内存中的数据存储

  • 避免全局列表data2存储过多数据,每爬取20-30页就将数据写入CSV并清空列表:
    if current_page_number % 20 == 0:
        df2 = pd.DataFrame(data2, columns=columns)
        df2.to_csv(f"scraped_profiles/scraped_profiles_page_{current_page_number}.csv", index=False)
        data2.clear()  # 清空列表释放内存
        import gc
        gc.collect()  # 强制触发垃圾回收
    
  • 爬取完每页后,主动清空元素引用:
    # 爬取完一页后执行
    profile_elements = None
    gc.collect()
    

3. 启用无头浏览器模式

Chrome的UI渲染会占用大量内存,启用无头模式可显著降低内存消耗:

chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument('--disable-dev-shm-usage')  # 解决Linux环境下的内存限制问题

4. 绕过前端渲染,直接调用API接口

  • 打开浏览器开发者工具(F12),切换到「网络」标签,点击「下一页」时观察XHR/fetch请求,找到返回分页数据的API接口
  • 使用requests库直接请求该接口,无需加载整个页面,内存占用会大幅降低
  • 注意复制登录后的Cookie、Authorization等身份验证参数到请求头中,确保接口请求有效

5. 优化分页跳转逻辑

  • 检查页面是否存在可直接输入页码的输入框,若有则直接输入目标页码并提交,替代多次点击「下一页」:
    from selenium.webdriver.common.keys import Keys
    
    # 假设存在页码输入框,需根据实际页面调整选择器
    page_input = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input.pagination-page-input")))
    page_input.clear()
    page_input.send_keys(target_page_number)
    page_input.send_keys(Keys.ENTER)
    
  • 检查URL是否支持分页参数(如?page=130),若支持则直接拼接URL跳转:
    driver.get(f"https://matches.marathimatrimony.com/listing?page={target_page_number}")
    

内容的提问来源于stack exchange,提问作者Prathamesh Sawant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 21:22:02