使用Python的BeautifulSoup与requests爬取网页无结果求助
Hey there! Let's sort out why your code isn't pulling the Confirmed case count (59,805) from covid19india.org. Here's the breakdown of the issue and fixes:
核心问题分析
The biggest reason your code returns nothing is dynamic content rendering. Modern sites like this use JavaScript to load and update real-time data. When you use requests.get(), you only get the initial static HTML—none of the JavaScript-rendered elements (like the level-item divs with case numbers) are present in that response. That's why your loop never finds any matching elements.
Even if we assume the HTML snippet you shared is what you're receiving, your class matching is fragile: using attrs={'class':'level-item is-cherry fadeInUp'} does an exact string match. If the site adjusts class order or adds extra whitespace, this will fail.
解决方案1:用Selenium处理动态内容(最可靠)
To grab dynamically loaded data, you need to simulate a browser to let JavaScript run first. Here's how to do it with Selenium:
First, install Selenium and download a browser driver (like ChromeDriver, matching your Chrome version):
pip install selenium
Then use this code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd case_type = [] count = [] # Initialize Chrome driver (ensure ChromeDriver is in your system PATH) driver = webdriver.Chrome() driver.get('https://www.covid19india.org/') try: # Wait up to 10 seconds for the case elements to load wait = WebDriverWait(driver, 10) level_items = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, '.level-item'))) for item in level_items: # Extract case category (Confirmed/Active/Recovered) case_elem = item.find_element(By.TAG_NAME, 'h5') case_type.append(case_elem.text.strip()) # Extract case count count_elem = item.find_element(By.TAG_NAME, 'h1') count.append(count_elem.text.strip()) finally: driver.quit() # Always close the browser when done df = pd.DataFrame({'Case Type': case_type, 'Count': count}) print(df)
This code waits for the page to fully load, then extracts all case data cleanly.
解决方案2:修正BeautifulSoup选择器(仅适用于静态内容)
If the site ever returns static data (unlikely now, but just in case), you can fix your original code with more robust selectors:
import requests from bs4 import BeautifulSoup as bs import pandas as pd case_type = [] count = [] url = requests.get('https://www.covid19india.org/') url.raise_for_status() # Catch request errors (e.g., 404, 500) soup = bs(url.content, 'html.parser') # Use CSS selectors to match divs with all required classes (order doesn't matter) for a in soup.select('div.level-item.is-cherry.fadeInUp'): b = a.find('h1') c = a.find('h5') if b and c: # Avoid errors if elements are missing case_type.append(c.text.strip()) count.append(b.text.strip()) df = pd.DataFrame({'Case Type': case_type, 'Count': count}) print(df)
Key improvements here:
- CSS selectors are more flexible than exact class string matches
url.raise_for_status()checks if your request was successful- Adding
if b and cprevents crashes if elements are missing .strip()cleans up extra spaces in the count text
内容的提问来源于stack exchange,提问作者Blue Bird

