如何从远程网页获取Img src?现有代码无输出求解决方案
Hey there! Let’s break down why you’re not getting any image srcs from that URL and walk through actionable fixes step by step.
1. 页面可能是JavaScript动态渲染的
Many modern sites load content dynamically with JavaScript—simple requests calls only fetch the initial static HTML, which won’t include images loaded later by scripts. For this, you’ll need a tool that mimics a real browser to fully render the page.
Here’s a working example with Selenium (make sure you have the matching ChromeDriver installed for your browser version):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize browser driver = webdriver.Chrome() driver.get("https://www.vfmii.com/exc/aspquery?command=invoke&ipid=HL26423&ids=42337&RM=N") # Wait for images to load (explicit wait is more reliable than implicit) try: WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.TAG_NAME, "img")) ) except: print("Timed out waiting for images to load") # Extract all image srcs images = driver.find_elements(By.TAG_NAME, "img") for img in images: src = img.get_attribute("src") if src: # Fix relative paths if needed if not src.startswith(("http://", "https://")): src = f"https://www.vfmii.com{src}" print(src) driver.quit()
2. Anti-scraping measures are blocking your request
Sites often block requests without proper headers (like a valid User-Agent) or cookies. Try adding realistic headers to your requests call first:
import requests from bs4 import BeautifulSoup headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://www.vfmii.com/", "Accept-Language": "en-US,en;q=0.9" } response = requests.get( "https://www.vfmii.com/exc/aspquery?command=invoke&ipid=HL26423&ids=42337&RM=N", headers=headers ) # First check if you're getting a valid page print(f"Status Code: {response.status_code}") print("First 500 chars of response:\n", response.text[:500]) # If status code is 200, parse for images if response.status_code == 200: soup = BeautifulSoup(response.text, "html.parser") images = soup.find_all("img") for img in images: src = img.get("src") if src: if not src.startswith(("http://", "https://")): src = f"https://www.vfmii.com{src}" print(src) else: print("Request blocked or page not found")
3. Images might be inside an iframe
Some sites load content inside iframes. If that’s the case, you’ll need to switch to the iframe context first (using Selenium):
# After loading the page with Selenium try: # Locate the iframe (use ID, name, or XPath if needed) iframe = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "your-iframe-id")) ) driver.switch_to.frame(iframe) # Now extract images from the iframe images = driver.find_elements(By.TAG_NAME, "img") # ... process images as before ... # Switch back to the main page when done driver.switch_to.default_content() except: print("No iframe found or timed out")
4. Double-check your selector logic
If you’re using custom selectors instead of img tags, make sure they’re targeting the right elements. Use browser dev tools (F12) to inspect the page and confirm the image elements’ structure.
内容的提问来源于stack exchange,提问作者Basit ali

