Python Selenium如何在正确位置添加换行符实现爬取内容按链接分组输出
问题原因
你当前的代码逻辑是每爬取到一个名字就立即执行一次文件写入操作,且每次写入都追加了换行符,才会出现每个名字单独占一行的错误结果。
修正方案
调整写入逻辑:处理单个链接时,先把该链接下爬取到的所有名字暂存到临时列表中,等该链接的所有内容抓取完成后,再把列表里的名字用, 拼接为单行字符串,统一写入文件后再加换行符。
修改后代码
from selenium import webdriver from selenium.webdriver.common.keys import Keys import time from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pickle as pkl import shutil, json, os, re from sys import platform from time import sleep from selenium import webdriver from selenium.common.exceptions import \ NoSuchElementException, ElementClickInterceptedException, \ StaleElementReferenceException from selenium.webdriver.common.keys import Keys PATH = "C:\Program Files (x86)\chromedriver.exe" driver = webdriver.Chrome(PATH) # 改用with语句打开link.txt,避免文件句柄泄漏 with open("link.txt", "r", encoding="utf-8") as one: for two in one.readlines(): driver.get(two.strip()) # strip()去除链接首尾的换行、空白符,避免请求报错 name_list = [] for element in driver.find_elements_by_class_name('sc-1b6bgon-7'): name_list.append(element.text) # 单个链接的所有名字收集完成后再统一写入 with open("names.txt", "a", encoding="utf-8") as testtxt: testtxt.write(", ".join(name_list) + "\n")
额外优化说明
- 新增的
strip()方法可以去除读取到的链接首尾的换行、空格等无效字符,避免请求链接时出错 - 所有文件打开操作都指定了
encoding="utf-8",避免读写时出现中文乱码问题 - 用with语句管理文件打开操作,无需手动关闭文件,避免文件句柄泄漏
内容的提问来源于stack exchange,提问作者zackchess3
相关产品推荐
相关产品推荐

