You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何从现有数组为for循环的每次迭代分配列标题

解决Python循环中为列分配标题并生成结构化表格的问题

问题分析

你的原代码存在两个核心问题:

  1. 全局计数器counter会跨URL累计,导致第一个URL取满5个元素后,后续URL无法收集数据
  2. 所有数据存入一维列表,无法区分不同URL对应的列数据

要实现目标,需要先按URL单独收集每列的5条数据(不足补None),再用pandas将数据整理为带指定列标题的表格。

解决方案代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.common.exceptions import ElementNotVisibleException, NoSuchElementException

# 存储每列的5条数据
column_data = []
col_titles = ['30024', '30033', '30038']
urls = [
    'https://www.example.com/page1',
    'https://www.example.com/page2',
    'https://www.example.com/page3'
]

for url in urls:
    driver = webdriver.Chrome()  # 根据你的实际浏览器驱动调整
    driver.get(url)
    current_col = []
    counter = 1  # 每个URL单独计数,避免跨URL干扰
    
    try:
        h2s = driver.find_elements(By.TAG_NAME, 'h2')
        for h2 in h2s:
            if counter <= 5:
                current_col.append(h2.get_attribute("innerText"))
                counter += 1
            else:
                break  # 收集够5个就停止遍历
        
        # 如果收集到的元素不足5个,补None填充到5条
        while len(current_col) < 5:
            current_col.append("None")
    
    except (ElementNotVisibleException, NoSuchElementException):
        # 发生异常时,直接填充5个None
        current_col = ["None"] * 5
    
    column_data.append(current_col)
    driver.close()

# 转置数据并创建DataFrame,指定列标题
df = pd.DataFrame(zip(*column_data), columns=col_titles)
# 输出不带索引的结构化表格
print(df.to_string(index=False))

关键改动说明

  1. 按列收集数据:每个URL对应生成一个长度固定为5的列表current_col,确保每列数据行数一致
  2. 独立计数器:每个URL内部重置counter,避免跨URL的计数干扰
  3. 补全数据:收集到的h2元素不足5个时,自动补None;发生异常时直接填充5个None
  4. 生成结构化表格:用zip(*column_data)将列数据转置为行数据,再通过pandas.DataFrame指定列标题,最终输出符合要求的表格

内容的提问来源于stack exchange,提问作者VRapport

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 04:35:50