使用Python Requests库获取Indeed职位页面HTML内容异常的问题求助
Ah, I’ve run into this exact issue with Indeed before—super frustrating when you’re trying to scrape job data and all you get is a bunch of JS boilerplate instead of actual listings. Let’s break down why this is happening and how to fix it:
1. Invalid URL Syntax
First off, your URL has an HTML-encoded & instead of a plain & between the query parameters. In Python, when you pass a URL string to requests.get(), you need to use the actual & character to separate parameters. Using & will make the server interpret the location parameter incorrectly, which might be part of why you're not getting the right content. The correct URL should be:https://uk.indeed.com/jobs?q=python&l=York
2. Dynamic Content Loading with JavaScript
Indeed doesn’t serve all job content directly in the initial HTML response. Instead, it uses JavaScript to fetch and render job listings after the page loads. The requests library only grabs the raw HTML sent by the server, which includes the scripts needed to load the actual content—but not the content itself (hence all those polyfill scripts you’re seeing).
3. Anti-Scraping Measures
Indeed actively blocks bot traffic. When you send a request with requests, you’re not sending the same headers as a real web browser. Sites like Indeed check for things like the User-Agent header to determine if the request is coming from a human or a bot. If they detect a bot, they might serve a stripped-down page or even a captcha instead of the job listings.
Fixes to Try
Option 1: Fix the URL and Add Realistic Headers
First, correct the URL, then mimic a real browser by adding headers like User-Agent to your request. Here’s how to modify your code:
from bs4 import BeautifulSoup import requests # Corrected URL with plain & instead of & url = 'https://uk.indeed.com/jobs?q=python&l=York' # Mimic a real Chrome browser's user agent headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } html_text = requests.get(url, headers=headers).text soup = BeautifulSoup(html_text, 'html.parser') # Now try searching for job elements, e.g., soup.find_all('div', class_='job_seen_beacon')
Note: You can update the User-Agent to match your current browser’s (search "what's my user agent" online to get yours).
Option 2: Use a Headless Browser to Render JavaScript
If adding headers doesn’t work (Indeed might still detect bot traffic), use a tool that can render the page like a real browser. Selenium or Playwright are great for this. Here’s a quick example with Selenium:
First, install Selenium and a webdriver (like ChromeDriver):
pip install selenium
Then modify your code:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options options = Options() options.add_argument('--headless=new') # Run without opening a visible browser window options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=options) driver.get('https://uk.indeed.com/jobs?q=python&l=York') # Wait a second to let the JS load the job listings (adjust timing if needed) driver.implicitly_wait(1) html_text = driver.page_source soup = BeautifulSoup(html_text, 'html.parser') # Now you should be able to find the job listings driver.quit()
Important: Make sure to respect Indeed’s Terms of Service when scraping. Avoid making too many requests too quickly, as this can get you blocked.
内容的提问来源于stack exchange,提问作者Quack

