Python循环中df.to_csv后台运行停止写入文件问题求助
问题
我正在使用Selenium遍历一个大型CSV文件,希望将结果追加到另一个CSV文件中,采用df.to_csv实现。起初运行正常,但当我切换到其他操作让程序后台运行时,CSV文件停止写入,不过循环仍在正常执行。当我回到Python程序界面并停留几秒后,CSV又恢复写入。
我查阅过相关解决方案,有用户建议使用.flush()方法,但该方法适用于将CSV文件存于变量中的场景,而我并未采用这种方式,因此找不到适用于追加CSV操作的flush方法。
代码示例:
while a < len(all_urls_snippet): last_last_height = 0 last_height = 0 now_height = 0 uniques = set() driver.get(all_urls_snippet[a]) time.sleep(3) now_height = driver.execute_script("return document.body.scrollHeight") while True: if last_last_height == last_height == now_height: break driver.execute_script("window.scrollBy(0, 400)") time.sleep(1.5) last_last_height = last_height last_height = now_height now_height = driver.execute_script("return document.body.scrollHeight") text = driver.find_elements("xpath", ".//div[contains(@class, 'status cursor-pointer focusable')]") for b in range(len(text)): comment_id = text[b].find_element("xpath", ".//div[contains(@class, 'status__wrapper space-y-4 status-public status-reply p-4')]") if comment_id.get_attribute("data-id") in uniques: continue else: uniques.add(comment_id.get_attribute("data-id")) like_button = comment_id.find_elements("xpath", ".//button[contains(@title, 'Like')]") for i in range(len(like_button)): if like_button[i].text == "": like_buttons.append(0) else: like_buttons.append(int(like_button[i].text)) # <<more code like this>> a = a + 1 columnsAll = [<<columns>>] df = pd.DataFrame(columnsAll) df = df.T df.to_csv('comments.csv', mode='a', index=False, header=False)
编辑补充:会不会是因为我在程序后台运行时同时运行了大型游戏导致的?不确定具体原因,无论如何,有没有办法强制程序每次都完成保存?
解决方案
强制立即写入磁盘的方法
直接用Python内置文件对象配合刷新操作,绕过pandas默认缓冲机制,确保每次写入都立即落地到磁盘:
import os # ... 你的循环逻辑 ... columnsAll = [<<columns>>] df = pd.DataFrame(columnsAll).T # 打开文件并执行写入 with open('comments.csv', 'a', newline='', encoding='utf-8') as f: df.to_csv(f, index=False, header=False) f.flush() # 将用户空间缓存刷新到内核 os.fsync(f.fileno()) # 强制内核把数据写入物理磁盘
这种方式不受程序前后台状态影响,数据不会卡在内存缓存里。
针对大型游戏场景的资源优化
如果同时运行大型游戏导致系统资源被抢占,可从以下几点调整:
- 切换无头浏览器模式:减少GPU和内存消耗,避免与游戏争夺资源:
from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument("--headless=new") # 新版无头模式,兼容性更强 chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options) - 替换固定sleep为显式等待:减少不必要的等待时间,提升程序运行效率:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 示例:等待目标元素加载完成再执行后续操作 WebDriverWait(driver, 10).until( EC.presence_of_element_located(("xpath", ".//div[contains(@class, 'status cursor-pointer focusable')]")) ) - 提升Python进程优先级:在系统中调高Python进程的优先级,避免被游戏进程抢占CPU。Windows可通过任务管理器设置,Linux/macOS使用
nice命令调整。
内容的提问来源于stack exchange,提问作者NewCoder1423
相关产品推荐
相关产品推荐

