Python批量提取多地址对应网页表格时仅返回单条结果求助
问题根因
- 遍历房产id的循环中,成功读取第一个地址的表格后执行了
break语句,直接终止了整个循环,仅能获取单个链接的结果 - 重复调用
pd.read_html(driver.page_source)会重复解析页面源码,效率低下,且未设置页面加载等待,部分页面表格未完全渲染就执行读取,会导致结果缺失 - 空
except语句吞掉了所有运行报错,无法定位具体出错的地址和环节 - 第一步获取mpr_id的循环写法冗余,写死9000的遍历上限存在索引越界风险
修正后代码
import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import pandas as pd endpoint = 'https://parser-external.geo.moveaws.com/suggest' params = { 'input': '', 'client_id': 'rdcV8', } headers = { 'Accept': 'application/json', 'Accept-Encoding': 'gzip, deflate, br', 'Accept-Language': 'en-US,en;q=0.9', 'Cache-Control': 'no-cache', 'Connection': 'keep-alive', 'DNT': '1', 'Host': 'parser-external.geo.moveaws.com', 'Origin': 'https://www.realtor.com', 'Pragma': 'no-cache', 'Referer': 'https://www.realtor.com/', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': 'Windows', 'Sec-Fetch-Dest': 'empty', 'Sec-Fetch-Mode': 'cors', 'Sec-Fetch-Site': 'cross-site', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/94.0.4606.61 Safari/537.36', } address = ['238 Lincoln St, Hahnville, LA 70057, USA', '101 Home Pl Ln, Hahnville, LA 70057, USA', '1250 Poydras St, New Orleans, LA 70113, USA', '1117 Broadway STE 401, Tacoma, WA 98402, USA', '2715 N Junett St, Tacoma, WA 98407, USA'] docs = [] for add in address: params['input'] = add try: r = requests.get(endpoint, params=params, headers=headers) r.raise_for_status() data = r.json() # 优化mpr_id提取逻辑 for item in data.get('autocomplete', []): if 'mpr_id' in item: document = 'M' + item['mpr_id'] docs.append(document) print(document) break except Exception as e: print(f"地址{add}获取mpr_id失败:{str(e)}") continue property_info_url = 'https://www.realtor.com/realestateandhomes-detail/' options = Options() options.headless = True # 新版selenium推荐使用Service指定chromedriver路径,可根据自身版本调整 # from selenium.webdriver.chrome.service import Service # s = Service('/Users/chromedriver/chromedriver') # driver = webdriver.Chrome(service=s, options=options) driver = webdriver.Chrome(executable_path='/Users/chromedriver/chromedriver',options=options) # 全局等待最长10秒,等待元素加载 wait = WebDriverWait(driver, 10) # 可替换为上面拿到的docs数组 test = ['M8021148490', 'M7965659502', 'M1205039451', 'M2499411953', 'M2907759924'] data={'first_table':[], 'second_table':[], 'third_table':[], 'fourth_table':[], 'fifth_table':[], 'sixth_table':[]} for prop_id in test: print(f"正在获取id{prop_id}的信息...") try: driver.get(property_info_url + prop_id) # 等待表格加载完成 wait.until(EC.presence_of_element_located((By.TAG_NAME, 'table'))) # 一次解析所有表格,避免重复调用 all_tables = pd.read_html(driver.page_source) if len(all_tables) >=6: data['first_table'].append(all_tables[0]) data['second_table'].append(all_tables[1]) data['third_table'].append(all_tables[2]) data['fourth_table'].append(all_tables[3]) data['fifth_table'].append(all_tables[4]) data['sixth_table'].append(all_tables[5]) else: print(f"id{prop_id}页面表格数量不足6个,仅获取到{len(all_tables)}个") except Exception as e: print(f"id{prop_id}获取失败:{str(e)}") continue driver.quit() print(data)
补充说明
- 提取到的每个表格都是pandas DataFrame格式,可根据需求合并后导出为csv、excel等文件
- 可根据实际页面表格结构调整判断逻辑,适配不同页面的表格数量差异
内容的提问来源于stack exchange,提问作者Stackbeans
相关产品推荐
相关产品推荐

