网页爬取问题:无<tr>标签仅含<div>的表格中.table-row类数据无法获取
.table-row Element Issue in Web Scraping Hey there! Let's break down why you're able to grab .table-header and .practiceDataTable elements but not .table-row ones—this is a super common scenario when dealing with modern websites.
The Likely Culprit: Dynamic Content Loading
When you use requests.get(), you're only fetching the initial static HTML of the page. Many modern sites (like the NASCAR results page you're targeting) load additional content after the initial page loads using JavaScript.
Chances are, the .table-row elements are being rendered dynamically by JS long after the initial HTML is sent to your browser. That's why they don't show up in the response from requests—they weren't there when the page first loaded! The .table-header and .practiceDataTable are part of the static initial markup, so they come through just fine.
How to Fix It
Here are two reliable solutions to get those .table-row elements:
1. Simulate a Browser with Selenium
Selenium launches a real browser (or headless browser) that executes JavaScript just like a human would, so it waits for dynamic content to load before you scrape it. Here's a working example:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import bs4 # Initialize Chrome (make sure you have chromedriver installed and in your PATH) driver = webdriver.Chrome() try: # Load the target page driver.get('https://www.nascar.com/results/race_center/2018/m...') # Wait up to 10 seconds for the first .table-row element to appear WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'table-row')) ) # Grab the fully rendered page source page_source = driver.page_source soup = bs4.BeautifulSoup(page_source, 'html.parser') # Now you can access the .table-row elements! table_rows = soup.find_all(class_='table-row') for row in table_rows: print(row.get_text(strip=True)) finally: # Always close the browser when done driver.quit()
2. Directly Call the Backend API (More Efficient)
Instead of simulating a browser, you can find the API endpoint that the site uses to load the table data. Here's how:
- Open your browser's DevTools (press F12), go to the Network tab
- Refresh the page, filter requests by XHR/Fetch
- Look for requests that return JSON data containing the race results (the URL might look like
/api/race-results/2018/...) - Copy that API URL, then use
requeststo fetch the JSON directly:
import requests api_url = "INSERT_THE_API_URL_YOU_FOUND_HERE" res = requests.get(api_url) data = res.json() # Parse the JSON data directly (no BeautifulSoup needed!) for row in data['results']: # Adjust the key based on the actual JSON structure print(row['driver_name'], row['position'], row['points'])
Quick Debugging Step
To confirm it's a dynamic loading issue, save the response from requests.get() to a file and open it in a browser:
res = requests.get('https://www.nascar.com/results/race_center/2018/m...') with open('page.html', 'w', encoding='utf-8') as f: f.write(res.text)
If you open page.html and don't see the .table-row elements, that confirms they're being added dynamically by JavaScript.
内容的提问来源于stack exchange,提问作者sbiondio

