You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python循环中df.to_csv后台运行停止写入文件问题求助

问题

我正在使用Selenium遍历一个大型CSV文件,希望将结果追加到另一个CSV文件中,采用df.to_csv实现。起初运行正常,但当我切换到其他操作让程序后台运行时,CSV文件停止写入,不过循环仍在正常执行。当我回到Python程序界面并停留几秒后,CSV又恢复写入。

我查阅过相关解决方案,有用户建议使用.flush()方法,但该方法适用于将CSV文件存于变量中的场景,而我并未采用这种方式,因此找不到适用于追加CSV操作的flush方法。

代码示例:

while a < len(all_urls_snippet):
    last_last_height = 0
    last_height = 0
    now_height = 0

    uniques = set()
    driver.get(all_urls_snippet[a])
    time.sleep(3)
    now_height = driver.execute_script("return document.body.scrollHeight")
    while True:
        if last_last_height == last_height == now_height:
            break
                    
        driver.execute_script("window.scrollBy(0, 400)")
        time.sleep(1.5)
        last_last_height = last_height
        last_height = now_height
        now_height = driver.execute_script("return document.body.scrollHeight")
        text = driver.find_elements("xpath", ".//div[contains(@class, 'status cursor-pointer focusable')]")

        for b in range(len(text)):
            comment_id = text[b].find_element("xpath", ".//div[contains(@class, 'status__wrapper space-y-4 status-public status-reply p-4')]")
            if comment_id.get_attribute("data-id") in uniques:
                continue
            else:
                uniques.add(comment_id.get_attribute("data-id"))
                like_button = comment_id.find_elements("xpath", ".//button[contains(@title, 'Like')]")
                for i in range(len(like_button)):
                    if like_button[i].text == "":
                        like_buttons.append(0)
                    else:
                        like_buttons.append(int(like_button[i].text))
            # <<more code like this>>
    
    a = a + 1

    columnsAll = [<<columns>>]
    df = pd.DataFrame(columnsAll)
    df = df.T
    df.to_csv('comments.csv', mode='a', index=False, header=False)

编辑补充:会不会是因为我在程序后台运行时同时运行了大型游戏导致的?不确定具体原因,无论如何,有没有办法强制程序每次都完成保存?

解决方案

强制立即写入磁盘的方法

直接用Python内置文件对象配合刷新操作,绕过pandas默认缓冲机制,确保每次写入都立即落地到磁盘:

import os
# ... 你的循环逻辑 ...

columnsAll = [<<columns>>]
df = pd.DataFrame(columnsAll).T
# 打开文件并执行写入
with open('comments.csv', 'a', newline='', encoding='utf-8') as f:
    df.to_csv(f, index=False, header=False)
    f.flush()  # 将用户空间缓存刷新到内核
    os.fsync(f.fileno())  # 强制内核把数据写入物理磁盘

这种方式不受程序前后台状态影响,数据不会卡在内存缓存里。

针对大型游戏场景的资源优化

如果同时运行大型游戏导致系统资源被抢占,可从以下几点调整:

  • 切换无头浏览器模式:减少GPU和内存消耗,避免与游戏争夺资源:
    from selenium.webdriver.chrome.options import Options
    
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")  # 新版无头模式,兼容性更强
    chrome_options.add_argument("--disable-gpu")
    driver = webdriver.Chrome(options=chrome_options)
    
  • 替换固定sleep为显式等待:减少不必要的等待时间,提升程序运行效率:
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # 示例:等待目标元素加载完成再执行后续操作
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located(("xpath", ".//div[contains(@class, 'status cursor-pointer focusable')]"))
    )
    
  • 提升Python进程优先级:在系统中调高Python进程的优先级,避免被游戏进程抢占CPU。Windows可通过任务管理器设置,Linux/macOS使用nice命令调整。

内容的提问来源于stack exchange,提问作者NewCoder1423

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 17:47:29