You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用pandas+Selenium写入xlsx所有行重复最后结果如何解决

问题根源

你代码出现所有行重复最后一次查询结果的核心原因是pandas.DataFrame.assign()的用法错误:

  • assign方法是对整个DataFrame的指定列做全局赋值,你每次循环调用df1.assign(type = info1...)都会把type到type5这几列的所有行都修改为当前循环拿到的查询结果,最后一次循环的结果会覆盖掉之前所有次的赋值,最终所有行的内容都和最后一行一致。

另外还有几个会影响查询结果正确性的隐性问题:

  • 直接遍历df.info会和pandas内置的info()方法产生命名冲突,应该显式指定取df['info']列遍历
  • 每次完成搜索后没有清空搜索框,下一次输入的关键词会和上一次的内容拼接,导致查询内容和预期不符
  • 隐式等待只需要全局设置一次,不需要放在循环内重复定义
修复后的完整代码
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.firefox.firefox_binary import FirefoxBinary
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import pandas as pd
import openpyxl
import numpy as np
import time
import sys
import os
import unittest
import csv
import re
import easygui
os.environ['MOZ_HEADLESS'] = '1'

# 选择Excel文件
chemin = easygui.fileopenbox(msg=None, title='Selectionner votre fichier Excel', default='*.xlsx', filetypes='', multiple=False)
# 初始化浏览器
driver = webdriver.Firefox() 
# 全局设置隐式等待
val = 30
driver.implicitly_wait(val)
print("connexion au site")
print("démarage du process")
data = pd.read_excel(chemin)
df1 = pd.DataFrame(data)

# 初始化列表存储每行对应的查询结果
type_list = []
type2_list = []
type3_list = []
type4_list = []
type5_list = []

# 遍历info列的每行内容
for num in df1['info']:    
    print(num)
    # 等待搜索框加载
    element = WebDriverWait(driver, 30).until(EC.presence_of_element_located((By.ID, "search")))
    elem = driver.find_element_by_id("search")
    # 清空搜索框后输入内容
    elem.clear()
    elem.send_keys(str(num))
    elem.send_keys(Keys.RETURN)
    # 隐藏加载指示器
    driver.execute_script("document.getElementById('waiting-indicator').style.display = 'none';")
    # 等待结果加载
    WebDriverWait(driver, 30).until(EC.presence_of_element_located((By.XPATH, "/html/body/div[1]/div[5]/div[1]/div/div[2]/div/div[1]/div/p")))

    print("récupération des datas")
    info1 = driver.find_element_by_xpath('/html/body/div[1]/div[5]/div[1]/div/div[2]/div/div[1]/div/p').text
    info2 = driver.find_element_by_xpath('/html/body/div[1]/div[5]/div[1]/div/div[2]/div/div[2]/div[1]/p').text
    info3 = driver.find_element_by_xpath('/html/body/div[1]/div[5]/div[1]/div/div[2]/div/div[3]/div[1]/p').text
    info4 = driver.find_element_by_xpath('/html/body/div[1]/div[5]/div[1]/div/div[2]/div/div[3]/div[2]/p').text
    info5 = driver.find_element_by_xpath('/html/body/div[1]/div[5]/div[1]/div/div[2]/div/div[3]/div[3]/p').text
    
    # 将当前行的查询结果存入对应列表
    type_list.append(info1)
    type2_list.append(info2)
    type3_list.append(info3)
    type4_list.append(info4)
    type5_list.append(info5)
    print("next")

# 循环结束后一次性将结果列表赋值到DataFrame的对应列
df1['type'] = type_list
df1['type2'] = type2_list
df1['type3'] = type3_list
df1['type4'] = type4_list
df1['type5'] = type5_list

print("opération terminé")
df1.to_excel(chemin+"out.xlsx", index = False, header=True)
driver.close()
关键修改说明
  • 循环前提前初始化5个空列表,每次循环仅把当前查询的结果追加到列表末尾,保证列表顺序和DataFrame行顺序一一对应
  • 每次输入搜索关键词前调用elem.clear()清空搜索框,避免关键词叠加导致的查询错误
  • 移除循环内重复的隐式等待定义,全局只设置一次
  • 循环全部结束后,把存储了所有行结果的列表一次性赋值给DataFrame的对应列,替代原逻辑里每次全局覆盖的错误写法

内容的提问来源于stack exchange,提问作者Shilo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 09:45:04