BeautifulSoup无法获取iframe内部标签?原因排查及爬取方案咨询
1. Why isn't my code getting the iframe's inner HTML?
When you use requests.get() to fetch a webpage, you only receive the initial HTML content sent by the server. The <iframe> tag in this response is just a pointer to another URL (the src attribute) — it doesn’t include the iframe’s actual content. Browsers automatically load the iframe’s content by sending a separate HTTP request to that src URL, but requests doesn’t handle this automatically. That’s why your output only shows the iframe tag itself, not its inner HTML.
2. How to verify if content is loaded via JavaScript?
Here are a few practical ways to check:
- Compare page source vs. dev tools: Right-click the page and select "View Page Source". Search for the content you expect to see. If it’s present in dev tools but missing from the page source, it’s either loaded via JavaScript or via an iframe (like your case).
- Inspect the iframe’s
src: Your code already shows the iframe has asrcvalue of/item/sise_day.nhn?code=005930. Visit that full URL (https://finance.naver.com/item/sise_day.nhn?code=005930) directly in your browser — if the content you want is there, it’s a separate static page loaded via the iframe, not JavaScript. - Disable JavaScript: Turn off JavaScript in your browser settings (e.g., Chrome: Settings > Privacy and security > Site Settings > JavaScript). Reload the page — if the content disappears, it’s dynamically loaded with JavaScript. For your specific case, the iframe content should still load even without JS, since it’s a standalone page.
3. What libraries to use if content is JavaScript-loaded?
Option 1: Directly fetch the iframe's URL (simplest for your case)
Since the iframe points to a standalone URL, you can skip the main page and fetch that URL directly with requests:
import requests from bs4 import BeautifulSoup # Fetch the iframe's content directly url = 'https://finance.naver.com/item/sise_day.nhn?code=005930' response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Extract the daily price table daily_price_table = soup.find('table', class_='type2') print(daily_price_table.prettify())
Option 2: Use a headless browser for JS-rendered content
If the content is truly dynamic (loaded after the initial page load via JavaScript), you’ll need a tool that mimics a real browser. Popular Python options are Selenium and Playwright.
Example with Selenium:
First, install the required packages:
pip install selenium webdriver-manager
Then use this code:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager from bs4 import BeautifulSoup url = 'https://finance.naver.com/item/sise.nhn?code=005930' # Initialize Chrome browser (auto-installs ChromeDriver) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get(url) # Switch to the iframe by its title driver.switch_to.frame(driver.find_element(By.CSS_SELECTOR, 'iframe[title="일별 시세"]')) # Get the fully rendered HTML of the iframe iframe_html = driver.page_source soup = BeautifulSoup(iframe_html, 'html.parser') # Extract your desired data daily_price_table = soup.find('table', class_='type2') print(daily_price_table.prettify()) # Clean up: close the browser driver.quit()
内容的提问来源于stack exchange,提问作者Sambo Kim

