使用Python BeautifulSoup爬取OpenTable网站遇大量NA列问题求助
解决OpenTable爬取中NA列过多的问题
看起来你的代码逻辑框架没问题,但大量NA列出现主要是因为OpenTable网页的元素class名称已经更新,再加上部分元素的定位方式不够健壮。我帮你逐一排查并优化代码:
1. 修复星级评分(rating)的提取逻辑
你原来用div.all-stars.filled定位星级,但现在网站的星级元素class已经改成了allStars__filled,而且星级数值不再通过style属性存储,而是放在元素的aria-label里。修改后的代码:
rating_container = resto.find('div', class_='allStars') if rating_container: # 提取类似"4.5 stars"里的数值 rating_text = rating_container.get('aria-label') item['rating'] = float(rating_text.split()[0]) if rating_text else 'NA' else: item['rating'] = 'NA'
2. 修复评论数(reviews)的提取逻辑
原选择器span.star-rating-text--review-text已经失效,现在评论数藏在a.starRating__reviewLink标签里,调整提取方式:
reviews_link = resto.find('a', class_='starRating__reviewLink') if reviews_link: # 从"(123 reviews)"这类文本中提取数字 reviews_text = reviews_link.text.strip() item['reviews'] = int(re.search(r'\d+', reviews_text).group()) if reviews_text else 'NA' else: item['reviews'] = 'NA'
3. 修复菜系(cuisine)的提取逻辑
原classrest-row-meta--cuisine已更新,现在菜系和位置都在div.rest-row-meta下的span.rest-row-meta--text标签里,需要按顺序提取:
meta_spans = resto.find_all('span', class_='rest-row-meta--text') if len(meta_spans) >= 2: item['cuisine'] = meta_spans[0].text.strip() item['location'] = meta_spans[1].text.strip() else: item['cuisine'] = 'NA' item['location'] = meta_spans[0].text.strip() if meta_spans else 'NA'
4. 优化数据存储方式(更高效稳定)
你原来用data[i] = pd.Series(item)的方式效率较低,还容易出现索引问题,建议先把每个餐厅的信息存入列表,最后统一转成DataFrame:
def parse_html(html): items = [] soup = BeautifulSoup(html, 'lxml') for resto in soup.find_all('div', class_='rest-row-info'): item = {} # 提取餐厅名称 name_elem = resto.find('span', class_='rest-row-name-text') item['name'] = name_elem.text.strip() if name_elem else 'NA' # 提取预订数 booking = resto.find('div', class_='booking') item['bookings'] = re.search(r'\d+', booking.text).group() if booking else 'NA' # 插入上面修复的rating、reviews、cuisine/location代码 # 提取价格等级 pricing_elem = resto.find('div', class_='rest-row-pricing') if pricing_elem: price_symbols = pricing_elem.find_all('i', class_='rest-row-pricing__icon') item['price'] = len(price_symbols) if price_symbols else 'NA' else: item['price'] = 'NA' items.append(item) return pd.DataFrame(items)
5. 完善翻页逻辑(原来的代码注释掉了翻页)
你注释掉了翻页代码,而且原逻辑没判断最后一页的情况,容易陷入死循环。调整后:
from selenium.common.exceptions import NoSuchElementException restaurants = pd.DataFrame() driver = webdriver.Chrome(ChromeDriverManager().install()) url = "https://www.opentable.com/new-york-restaurant-listings" driver.get(url) while True: sleep(1) new_data = parse_html(driver.page_source) if new_data.empty: break restaurants = pd.concat([restaurants, new_data], ignore_index=True) print(f"已爬取{len(restaurants)}家餐厅") # 处理翻页:判断"Next"按钮是否存在且可点击 try: next_btn = driver.find_element_by_link_text('Next') # 如果按钮有disabled类,说明是最后一页 if 'disabled' in next_btn.get_attribute('class'): break next_btn.click() except NoSuchElementException: # 找不到Next按钮,退出循环 break driver.close() restaurants.to_csv('results.csv', index=False) print(restaurants)
这些修改应该能大幅减少NA列的出现,同时让爬取逻辑更健壮。另外,建议在实际运行时适当延长sleep时间,避免触发网站反爬机制。
内容的提问来源于stack exchange,提问作者M.monica
相关产品推荐
相关产品推荐

