You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取时如何定位表格精确Class?请求修复指定网站爬取代码

Fixing Table Scraping on Knoema

Your issue is that the table on this Knoema page is loaded dynamically using JavaScript. When you use requests.get(), you're only fetching the initial static HTML, which doesn’t include the table content yet. That’s why soup.find_all("table", class_="rank") returns an empty list—there’s no such table in the static response.

Step 1: How to Find the Exact Table Selector (Using Browser Dev Tools)

To confirm the table’s actual attributes once it’s loaded:

  • Open the URL in Chrome or Firefox.
  • Right-click anywhere on the table you want to scrape and select Inspect.
  • In the DevTools Elements tab, look for the <table> element wrapping the table content. For this specific page, the table has a class k-selectable and is nested inside a <div> with class k-grid-content.

Step 2: Solution Using Selenium (to Load Dynamic Content)

Since the table is rendered client-side, we need a tool that can execute JavaScript and wait for the table to load. Selenium is ideal for this. Here’s how to adjust your code:

First, install Selenium and a compatible webdriver (like ChromeDriver):

pip install selenium
# Download ChromeDriver from the official site (match your Chrome browser version)

Then use this updated code:

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize Chrome driver (ensure ChromeDriver is in your PATH or specify its path)
driver = webdriver.Chrome()
url = "https://knoema.com/atlas/topics/Tourism/Travel-and-Tourism-Total-Contribution-to-GDP/Contribution-of-travel-and-tourism-to-GDP"
driver.get(url)

# Wait up to 10 seconds for the table to load
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "table.k-selectable")))

# Get the fully rendered page source
soup = BeautifulSoup(driver.page_source, 'html.parser')

# Extract the target table
table_element = soup.find("table", class_="k-selectable")

# Convert table to pandas DataFrame
df = pd.read_html(str(table_element))[0]

# Clean up multi-row headers (optional but improves readability)
df = df.droplevel(0, axis=1)

print(df.head())

# Close the browser window
driver.quit()

Alternative: Use the Site’s API (Faster Option)

You can also check for an API endpoint that serves the table data directly:

  • Open DevTools > Network tab.
  • Refresh the page and look for XHR/fetch requests returning JSON data.
  • For Knoema, you’ll likely find a request with the raw table data, which you can parse directly with requests and convert to a DataFrame without needing Selenium. This is faster but requires inspecting the network traffic to locate the correct endpoint.

Why Your Original Code Failed

The requests.get() method only retrieves the initial HTML sent by the server. The table content is added later by JavaScript running in the browser, so it’s not present in the static response. That’s why your find_all call returned zero results.

内容的提问来源于stack exchange,提问作者Gambler Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:40:54