如何无需物理滚动抓取整页数据?Airtable表格全量爬取遇阻
问题:Airtable页面数据抓取仅获取前18行,如何加载全部2063行
我用以下代码抓取某Airtable网页数据,但只能获取前18行,无法加载全部2063行数据:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait import time driver = webdriver.Chrome() url= "https://airtable.com/shrqYt5kSqMzHV9R5/tbl8c8kanuNB6bPYr?backgroundColor=green&viewControls=on" driver.maximize_window() driver.get(url) time.sleep(5) content = driver.page_source.encode('utf-8').strip() soup = BeautifulSoup(content,"html.parser") table1 = soup.find("div",{"id":"table"}) driver.quit() company_names = [] for i in table1.find_all('div',class_="line-height-4 overflow-hidden truncate"): title = i.text if(title.startswith('http')): continue company_names.append(title) print(company_names) print(len(company_names))
我尝试了以下4种滚动网页的代码,但均无效:
方法1
ScrollNumber = 50 for i in range(1,ScrollNumber): driver.execute_script("window.scrollTo(1,5000)")#scrolling to said coordinates time.sleep(2) for i in range(1,ScrollNumber): driver.execute_script("window.scrollTo(5000,1)")#scrolling to said coordinates time.sleep(2) driver.close()
方法2
SCROLL_PAUSE_TIME = 0.5 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait to load page time.sleep(SCROLL_PAUSE_TIME) # Calculate new scroll height and compare with last scroll height new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height
方法3
elem = driver.find_element_by_tag_name("body") no_of_pagedowns = 38 while no_of_pagedowns: elem.send_keys(Keys.PAGE_DOWN) time.sleep(1) no_of_pagedowns-=1
方法4
elem = driver.find_element_by_tag_name('body') elem.send_keys(Keys.END)
解决方案
核心原因:Airtable表格的滚动容器不是body
Airtable的表格内容存放在独立的垂直滚动容器内,而非整个页面的body标签。之前的滚动操作都针对body,自然无法触发表格数据的加载,必须定位到正确的滚动容器进行操作。
正确的滚动实现代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time driver = webdriver.Chrome() url= "https://airtable.com/shrqYt5kSqMzHV9R5/tbl8c8kanuNB6bPYr?backgroundColor=green&viewControls=on" driver.maximize_window() driver.get(url) try: # 显式等待表格滚动容器加载完成(替代固定sleep,更稳定) scroll_container = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div[data-scrollable='vertical']")) ) # 循环滚动直到数据不再加载 last_row_count = 0 while True: # 获取当前已加载的行数 current_rows = driver.find_elements(By.CSS_SELECTOR, "div.line-height-4.overflow-hidden.truncate") current_row_count = len(current_rows) # 行数不再变化则停止滚动 if current_row_count == last_row_count: break last_row_count = current_row_count # 滚动到容器底部 driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight", scroll_container) time.sleep(1) # 等待数据加载,可根据网络情况调整 # 提取并处理数据 content = driver.page_source.encode('utf-8').strip() soup = BeautifulSoup(content,"html.parser") table1 = soup.find("div",{"id":"table"}) company_names = [] for i in table1.find_all('div',class_="line-height-4 overflow-hidden truncate"): title = i.text if(title.startswith('http')): continue company_names.append(title) print(company_names) print(f"总共获取到 {len(company_names)} 条数据") finally: driver.quit()
关键优化说明
- 定位正确容器:通过
div[data-scrollable='vertical']选择器定位Airtable的垂直滚动容器,这是平台的标准属性。 - 动态判断加载完成:对比滚动前后的行数,避免固定次数滚动导致的遗漏或无效操作。
- 显式等待:用WebDriverWait替代固定sleep,适配不同网络环境,提升稳定性。
备选高效方案:使用Airtable API
如果能获取Airtable的API密钥,直接调用官方接口获取数据会比浏览器模拟更高效、稳定,无需处理页面滚动逻辑。
内容的提问来源于stack exchange,提问作者random_name
相关产品推荐
相关产品推荐

