使用BeautifulSoup爬取Captcha图片链接失败问题求助
Hey Anna, I’ve run into this exact problem a few times—let’s walk through the most common reasons your code isn’t picking up that captcha image link, plus fixes for each scenario:
1. The captcha uses a relative URL instead of absolute
A lot of sites serve captchas with paths like /api/get-captcha instead of full URLs. When you print the img tag, you might only see that relative path, not the full link your browser resolves automatically.
Fix it with urllib.parse.urljoin:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin with requests.Session() as s: base_url = "https://myurl.com/" # Add a realistic user-agent to mimic a browser headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} r = s.get(base_url, headers=headers) soup = BeautifulSoup(r.content, "html.parser") for img in soup.find_all("img"): img_src = img.get("src") if img_src: # Convert relative path to full absolute URL full_img_url = urljoin(base_url, img_src) print(f"Full image URL: {full_img_url}") # Target the captcha specifically and download it if "captcha" in img_src.lower(): captcha_response = s.get(full_img_url, headers=headers) with open("captcha.png", "wb") as f: f.write(captcha_response.content)
2. The captcha is loaded dynamically with JavaScript
Requests only fetches the initial static HTML—if the captcha is rendered after the page loads via JS (like many modern sites do), BeautifulSoup won’t see it at all.
Use Selenium to mimic a real browser:
This lets you wait for JS to load the captcha before scraping:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() # Ensure ChromeDriver is installed and in your PATH driver.get("https://myurl.com/") # Wait up to 10 seconds for the captcha image to appear try: captcha_img = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//img[contains(@src, 'captcha')]")) ) captcha_src = captcha_img.get_attribute("src") print(f"Captcha URL: {captcha_src}") # Optional: Save the captcha directly via screenshot captcha_img.screenshot("captcha.png") finally: driver.quit()
3. The site blocks non-browser requests
Some sites flag requests without valid browser headers. Your current code uses requests’ default user-agent, which might get rejected or served a different version of the page.
Add realistic headers to your request:
As shown in the first code example, including a User-Agent (and even other headers like Accept-Language) makes your request look like it’s coming from a real browser, matching what you see in the dev tools.
4. The captcha is embedded as base64 data
Occasionally, sites serve captchas directly as base64-encoded data in the src attribute (looks like data:image/png;base64,iVBORw0KGgo...).
Extract and save the base64 data:
import base64 # Assuming you've located the captcha img tag captcha_src = img.get("src") if captcha_src.startswith("data:image"): # Split the base64 content from the metadata header img_data = captcha_src.split(",")[1] # Decode and save the image with open("captcha.png", "wb") as f: f.write(base64.b64decode(img_data))
Start by checking if the captcha URL is relative or if you’re missing headers—those are the quickest fixes. If that doesn’t work, dynamic JS loading is probably the culprit, and Selenium will resolve that.
内容的提问来源于stack exchange,提问作者Anna Plym

