使用Python Requests访问动态URL及MLB对阵数据爬取问题
Hey there, let's tackle your two related issues head-on—dynamic content not loading with Python Requests, and specifically scraping that MLB batter vs pitcher matchup page.
1. Fixing Dynamic Content When URL Has Request Parameters
If the page content shifts even when your URL includes the right parameters, here are the most common fixes to try:
Add realistic request headers
Many sites block requests that don’t look like they’re coming from a real browser. At minimum, include aUser-Agentheader to mimic a standard browser. Example:import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(your_target_url, headers=headers)You can grab other headers (like
Referer) from your browser’s DevTools Network tab to make the request even more authentic.Use
paramsinstead of hardcoding parameters in the URL
Manual URL concatenation often leads to encoding issues (like&instead of&). Let Requests handle parameter formatting automatically:params = { 'season': 2018, 'batting_team': 119, 'batter': 571771, 'pitching_team': 133, 'pitcher': 641941 } response = requests.get('http://mlb.mlb.com/stats/sortable_batter_vs_pitcher.jsp', params=params, headers=headers)This ensures all parameters are correctly encoded and avoids messy URL strings.
Check if content loads via JavaScript
Requests only fetches the initial static HTML. If the data loads later via AJAX/fetch calls, you have two solid options:- Find the underlying API: Open your browser’s DevTools, go to the Network tab, and filter for XHR/fetch requests. Look for calls that return the matchup data (usually in JSON format). Directly request this API endpoint instead of the main page—it’s faster and easier to parse.
- Simulate a browser: Use tools like Selenium or Playwright to load the page like a real user, which executes JavaScript. Example with Selenium:
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By driver = webdriver.Chrome() url = matchupURL(2018, 571771, 119, 641941, 133) driver.get(url) # Wait for the stats table to load (adjust the selector to match the page's actual elements) wait = WebDriverWait(driver, 10) stats_table = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.stats-table'))) # Grab the fully rendered HTML or extract data directly page_content = driver.page_source driver.quit()
2. Specific Fix for MLB Batter vs Pitcher Page
For that MLB stats page, here are extra tips tailored to the site:
- Fix URL formatting: Your
matchupURLfunction might generate URLs with&(an HTML entity for&). Replace those with plain&or use theparamsmethod above to avoid this issue entirely. - Hunt for hidden API endpoints: When loading the page in your browser, watch the Network tab for requests to endpoints like
/stats/json/...or similar. These often return structured JSON data that’s way easier to work with than scraping HTML. - No login needed (usually): MLB’s public stats pages don’t require authentication, so you shouldn’t need to handle login cookies. If you hit blocks, try copying cookies from your browser into the Requests
cookiesparameter to mimic your browser session.
内容的提问来源于stack exchange,提问作者Brock Lumbard

