You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何结合双循环/Zip数组将Selenium爬取数据写入CSV

Solution to Pair Selenium Elements and Write to CSV

Got it, let's get your CSV output sorted! The core issue here is pairing each company's link/name with its corresponding address—Python's zip() function is perfect for this, since it lets you iterate over both element lists in lockstep.

Here's how to adjust your code, with key fixes and improvements:

from selenium.webdriver.common.by import By
import csv

# Your existing element retrieval code
company_links_elements = driver.find_elements(By.XPATH, "//h3[@class='jss320 jss324 jss337 sc-gzOgki eucExu']/a")
company_address_elements = driver.find_elements(By.XPATH, "//strong[@class='dtm-search-listing-address']")

with open('links.csv', 'w', newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    # Write CSV header first (optional but recommended for clarity)
    writer.writerow(["company_name", "company_url", "company_address"])
    
    # Use zip() to iterate over both lists simultaneously
    for company_link, address_element in zip(company_links_elements, company_address_elements):
        company_url = company_link.get_attribute("href")
        # Fix: Use .text instead of get_attribute("text") for reliable element text extraction
        company_name = company_link.text.strip()  # .strip() cleans up extra whitespace/newlines
        company_address = address_element.text.strip()
        
        # Write all three values to a single CSV row
        writer.writerow([company_name, company_url, company_address])

driver.close()

Key Details to Note:

  • zip() Function: This pairs each item from company_links_elements with the item at the exact same index in company_address_elements—just make sure the two lists are in matching order (which they should be if scraped from the same page's consistent structure).
  • Element Text Fix: get_attribute("text") isn't the right way to grab an element's visible text. Use the .text property instead, and add .strip() to clean up messy whitespace that often comes with scraped text.
  • CSV Best Practices: Adding encoding='utf-8' ensures special characters (like accented letters or symbols) in company names/addresses save correctly, and newline='' prevents extra blank rows from appearing in Windows.

Handling Mismatched List Lengths:

If for some reason the two lists have different lengths (e.g., some companies don't display an address), use itertools.zip_longest to keep all entries even when one is missing:

from itertools import zip_longest
from selenium.webdriver.common.by import By
import csv

# ... (keep element retrieval code)

with open('links.csv', 'w', newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(["company_name", "company_url", "company_address"])
    
    # Use zip_longest to handle unequal lengths, fill missing values with "N/A"
    for company_link, address_element in zip_longest(company_links_elements, company_address_elements, fillvalue=None):
        company_url = company_link.get_attribute("href") if company_link else "N/A"
        company_name = company_link.text.strip() if company_link else "N/A"
        company_address = address_element.text.strip() if address_element else "N/A"
        
        writer.writerow([company_name, company_url, company_address])

This way, you won't lose any scraped data even if there's a mismatch in the number of elements.

内容的提问来源于stack exchange,提问作者Kyle Linden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:39:17