使用Selenium和Pandas爬取数据报错:值长度与索引不匹配
爬取物业交易数据时Pandas ValueError问题排查与解决
问题描述
爬取17/18页物业交易记录后,程序抛出如下错误,疑问是否为Pandas的问题,需求是将18页数据导出为CSV文件:
ValueError: Length of values (429) does not match length of index (18)
错误详情
Traceback (most recent call last): File "C:\Users\user\PycharmProjects\pythonProject2\venv\centranet_v1.chrome 20230219.py", line 98, in <module> df['Address'] = Address ~~^^^^^^^^^^^ File "C:\Users\user\PycharmProjects\pythonProject2\venv\Lib\site-packages\pandas\core\frame.py", line 3980, in __setitem__ self._set_item(key, value) File "C:\Users\user\PycharmProjects\pythonProject2\venv\Lib\site-packages\pandas\core\frame.py", line 4174, in _set_item value = self._sanitize_column(value) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\user\PycharmProjects\pythonProject2\venv\Lib\site-packages\pandas\core\frame.py", line 4915, in _sanitize_column com.require_length_match(value, self.index) File "C:\Users\user\PycharmProjects\pythonProject2\venv\Lib\site-packages\pandas\core\common.py", line 571, in require_length_match raise ValueError( ValueError: Length of values (429) does not match length of index (18) Process finished with exit code 1
原始代码
#!/usr/bin/env python # coding: utf-8 from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from time import sleep import datetime as dt import pandas as pd #open chrome web browser driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.maximize_window() driver.get("https://hk.centanet.com/findproperty/en/list/transaction") #find elements in web page wait = WebDriverWait(driver, 10) inputElem = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "input"))) inputElem.clear() inputElem.send_keys('Discovery Park') wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "btn-search"))).click() wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@class='btn-fiter']/span[contains(text(), 'Sold / Leased')]"))).click() wait.until(EC.element_to_be_clickable((By.XPATH, "//span[@class='el-radio__label']/span[contains(text(), 'Sold')]"))).click() click_next = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "btn-next"))) Date=Address=Price=Changes=Saleable_Area=[] def it(x): lst=[] for i in x: lst.append(i.text) return lst while True: content = driver.find_element(By.CSS_SELECTOR, "div[class*='bx--structured-list-td']") info_date = content.find_elements(By.CLASS_NAME,"info-date") Date = Date + it(info_date) info_address = content.find_elements(By.XPATH, "//div[@class='cv-structured-list-data bx--structured-list-td']/div[contains(text(), 'Discovery Park')]") Address = Address + it(info_address) tranprice = content.find_elements(By.CLASS_NAME,"tranPrice") Price = Price + it(tranprice) info_changes = content.find_elements(By.CLASS_NAME,"riseBox") Changes = Changes + it(info_changes) feet = content.find_elements(By.XPATH, "//div[@class='cv-structured-list-data bx--structured-list-td']/div[contains(text(), 'ft²')]") Saleable_Area = Saleable_Area + it(feet) if click_next.is_enabled(): click_next.click() else: break #break the while loop if click next is disable sleep(2) driver.quit() df = pd.DataFrame() df['Date'] = Date df['Address'] = Address df['Price'] = Price df['Changes%'] = Changes df['Saleable_Area'] = Saleable_Area df.to_csv("result2.csv", index=False)
问题分析与解决
这不是Pandas的问题,Pandas只是在校验数据一致性时抛出了错误,根源在爬虫逻辑的缺陷:
核心原因
- 元素定位范围错误:
- 用
content = driver.find_element(...)只获取了单个单元格,而非整页交易行。 - 后续XPath使用全局匹配(
//div[...]),没有限制在当前行/页面范围内,导致每次循环都会把页面上所有匹配的地址重复添加到Address列表,最终Address长度(429)远大于其他字段(比如Date长度18)。
- 用
- 字段对应关系缺失:没有按行遍历提取数据,无法保证每行的字段一一对应,导致各列表长度不一致。
修复方案
1. 重构爬虫循环,按行提取数据
将原来的按字段批量获取,改为遍历每一行,在每行内部定位字段,确保数据对应且长度一致:
Date=Address=Price=Changes=Saleable_Area=[] while True: # 定位当前页的所有交易行 rows = driver.find_elements(By.CSS_SELECTOR, "div.bx--structured-list-row") for row in rows: # 提取日期(每行必存在) date_elem = row.find_element(By.CLASS_NAME, "info-date") Date.append(date_elem.text) # 提取地址(在当前行范围内查找,XPath前加.) addr_elem = row.find_element(By.XPATH, ".//div[contains(text(), 'Discovery Park')]") Address.append(addr_elem.text) # 提取价格 price_elem = row.find_element(By.CLASS_NAME, "tranPrice") Price.append(price_elem.text) # 提取涨跌(处理可能缺失的情况) try: change_elem = row.find_element(By.CLASS_NAME, "riseBox") Changes.append(change_elem.text) except: Changes.append("") # 提取实用面积 area_elem = row.find_element(By.XPATH, ".//div[contains(text(), 'ft²')]") Saleable_Area.append(area_elem.text) # 翻页逻辑 if click_next.is_enabled(): click_next.click() sleep(2) else: break
2. 关键优化点
- 限制查找范围:XPath前添加
.,表示仅在当前row元素下查找,避免全局匹配重复数据。 - 异常处理:对可能缺失的字段(如
Changes)用try-except填充空值,保证所有列表长度一致。 - 按行提取:确保每行的所有字段对应添加,不会出现长度差异。
3. 验证数据一致性
在创建DataFrame前,先打印各列表长度,确认一致:
print(f"Date长度: {len(Date)}, Address长度: {len(Address)}, Price长度: {len(Price)}, Changes长度: {len(Changes)}, Saleable_Area长度: {len(Saleable_Area)}")
修复后,各字段列表长度一致,即可正常创建DataFrame并导出CSV文件。
内容的提问来源于stack exchange,提问作者YPT
相关产品推荐
相关产品推荐

