Python电商爬取:Telegram Bot监控Myntra商品信息失败求助
Hey there! Let's troubleshoot why your Myntra price/product name scraping is returning None while Flipkart works fine. Here are the key fixes and improvements you can make:
1. Fix the Typo in Request Headers
First off, you have a small but critical typo in your headers: the key should be User-Agent (singular, capitalized correctly), not user_agents. Myntra's server likely rejects requests with malformed headers, leading to incomplete or empty page content that BeautifulSoup can't parse properly.
2. Improve Platform Detection Logic
Your current way of checking if the URL is for Myntra (splitting the HTML content by . and looking for "myntra") is unreliable. Instead, just check directly against the original URL—it's simpler and more accurate.
3. Correct Element Selection & Error Handling
For Myntra, you need to make sure you're extracting the text content from the elements you find (using .get_text()), and add basic checks to handle cases where elements might not be found.
Here's the revised code incorporating these fixes:
import requests from bs4 import BeautifulSoup # Example Myntra product URL URL = 'https://www.myntra.com/sports-sandals/roadster/roadster-men-charcoal-grey-sports-sandals/9024251/buy' # Fixed headers with correct User-Agent key headers = { "User-Agent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.88 Safari/537.36' } page = requests.get(URL, headers=headers) # First verify the request was successful if page.status_code != 200: print(f"Request failed! Status code: {page.status_code}") else: soup = BeautifulSoup(page.content, "html.parser") # Detect platform directly from URL if 'flipkart.com' in URL: title = soup.find(class_="_35KyD6").get_text(strip=True) price = soup.find(class_="_1vC4OE _3qQ9m1").get_text(strip=True) print("Flipkart Product Details:") print(f"Name: {title}") print(f"Price: {price}") elif 'myntra.com' in URL: # Target the highlighted product name and price elements product_name = soup.find(class_="pdp-name") product_price = soup.find(class_="pdp-price") if product_name and product_price: print("Myntra Product Details:") print(f"Name: {product_name.get_text(strip=True)}") print(f"Price: {product_price.get_text(strip=True)}") else: print("Couldn't find Myntra product details. This might be because Myntra uses dynamic JavaScript to load content.")
If You Still Get None—Handle Dynamic Content
Myntra often uses JavaScript to dynamically render product details, which means requests (only fetches static HTML) won't see the elements. In that case, use Selenium to simulate a real browser that loads the full page:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options URL = 'https://www.myntra.com/sports-sandals/roadster/roadster-men-charcoal-grey-sports-sandals/9024251/buy' # Configure headless Chrome to run without a visible window chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.88 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(URL) # Wait for the page to load elements (implicit wait for simplicity) driver.implicitly_wait(10) # Extract the highlighted product name and price product_name = driver.find_element(By.CLASS_NAME, "pdp-name").text product_price = driver.find_element(By.CLASS_NAME, "pdp-price").text print("Myntra Product Details (via Selenium):") print(f"Name: {product_name}") print(f"Price: {product_price}") driver.quit()
内容的提问来源于stack exchange,提问作者Anand Yeole

