使用BeautifulSoup监控Northvolt职位时HTML获取异常求助
Hey there! Let's break down why your requests + BeautifulSoup setup isn't pulling in the job listings you're looking for, and how to fix it.
Common Causes & Solutions
1. Dynamic Content Loading (Most Likely Culprit)
Modern websites like Northvolt often load content asynchronously with JavaScript after the initial HTML page loads. When you use requests.get(), you're only grabbing the raw, unrendered HTML sent by the server—this doesn't include the job listings, which are loaded later via JavaScript calls to a backend API.
If you check the "View Page Source" of the career page (not the browser inspector), you'll probably see empty container elements where jobs should appear, but no actual job text.
Fix 1: Use a Headless Browser to Render JavaScript
Tools like Selenium or Playwright simulate a real browser, executing all JavaScript and loading the full rendered page. Here's a quick Selenium example tailored to your use case:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from bs4 import BeautifulSoup # Set up headless Chrome to run without a visible window chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) driver.get("https://northvolt.com/career") try: # Wait up to 10 seconds for job elements to load (adjust the selector to match the page's actual elements) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "[data-test='job-card']")) ) # Grab the fully rendered page source page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # Check for your target job if "Energy coordinator" in soup.get_text(): print("✅ Energy coordinator position is open!") # Add your email notification logic here else: print("❌ Energy coordinator position is not open yet.") finally: driver.quit()
Fix 2: Directly Call the Job API (More Efficient)
Instead of rendering the whole page, you can find the backend API that feeds the job listings and request data directly. Here's how:
- Open your browser's DevTools (F12) → Go to the Network tab → Refresh the career page.
- Look for requests labeled "XHR" or "Fetch" that return job data (look for URLs containing "jobs", "career", or similar keywords).
- Copy that API URL and use
requeststo fetch the JSON data directly:
import requests # Replace this with the actual API URL you found in DevTools JOB_API_URL = "https://example-northvolt-api.com/career/jobs" response = requests.get(JOB_API_URL) response.raise_for_status() # Raise error if request fails job_data = response.json() # Check each job title for your target keyword for job in job_data: if "Energy coordinator" in job.get("title", "").lower(): print("✅ Energy coordinator position is open!") # Trigger your email notification here break else: print("❌ Energy coordinator position is not open yet.")
2. Anti-Scraping Measures
Northvolt might be blocking non-browser requests. To mimic a real browser, add headers to your requests calls (though this won't fix dynamic content issues):
headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } result = requests.get("https://northvolt.com/career", headers=headers)
Final Tips
- When using automated scrapers, respect the site's
robots.txtfile (checkhttps://northvolt.com/robots.txt) and limit request frequency to avoid getting blocked. - For daily checks, wrap your code in a scheduler (like the
schedulelibrary or cron jobs on Linux/macOS) to run it automatically.
内容的提问来源于stack exchange,提问作者Svenismyfriend

