如何结合双循环/Zip数组将Selenium爬取数据写入CSV
Solution to Pair Selenium Elements and Write to CSV
Got it, let's get your CSV output sorted! The core issue here is pairing each company's link/name with its corresponding address—Python's zip() function is perfect for this, since it lets you iterate over both element lists in lockstep.
Here's how to adjust your code, with key fixes and improvements:
from selenium.webdriver.common.by import By import csv # Your existing element retrieval code company_links_elements = driver.find_elements(By.XPATH, "//h3[@class='jss320 jss324 jss337 sc-gzOgki eucExu']/a") company_address_elements = driver.find_elements(By.XPATH, "//strong[@class='dtm-search-listing-address']") with open('links.csv', 'w', newline='', encoding='utf-8') as file: writer = csv.writer(file) # Write CSV header first (optional but recommended for clarity) writer.writerow(["company_name", "company_url", "company_address"]) # Use zip() to iterate over both lists simultaneously for company_link, address_element in zip(company_links_elements, company_address_elements): company_url = company_link.get_attribute("href") # Fix: Use .text instead of get_attribute("text") for reliable element text extraction company_name = company_link.text.strip() # .strip() cleans up extra whitespace/newlines company_address = address_element.text.strip() # Write all three values to a single CSV row writer.writerow([company_name, company_url, company_address]) driver.close()
Key Details to Note:
zip()Function: This pairs each item fromcompany_links_elementswith the item at the exact same index incompany_address_elements—just make sure the two lists are in matching order (which they should be if scraped from the same page's consistent structure).- Element Text Fix:
get_attribute("text")isn't the right way to grab an element's visible text. Use the.textproperty instead, and add.strip()to clean up messy whitespace that often comes with scraped text. - CSV Best Practices: Adding
encoding='utf-8'ensures special characters (like accented letters or symbols) in company names/addresses save correctly, andnewline=''prevents extra blank rows from appearing in Windows.
Handling Mismatched List Lengths:
If for some reason the two lists have different lengths (e.g., some companies don't display an address), use itertools.zip_longest to keep all entries even when one is missing:
from itertools import zip_longest from selenium.webdriver.common.by import By import csv # ... (keep element retrieval code) with open('links.csv', 'w', newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(["company_name", "company_url", "company_address"]) # Use zip_longest to handle unequal lengths, fill missing values with "N/A" for company_link, address_element in zip_longest(company_links_elements, company_address_elements, fillvalue=None): company_url = company_link.get_attribute("href") if company_link else "N/A" company_name = company_link.text.strip() if company_link else "N/A" company_address = address_element.text.strip() if address_element else "N/A" writer.writerow([company_name, company_url, company_address])
This way, you won't lose any scraped data even if there's a mismatch in the number of elements.
内容的提问来源于stack exchange,提问作者Kyle Linden
相关产品推荐
相关产品推荐

