Python新手如何抓取Amazon.jobs页面的所有Read More链接
Scraping Amazon.jobs Read More Links for Beginners
Hey there! Let's walk through how to grab all those "Read More" links from the Amazon.jobs search page you shared. Since you're a beginner, I'll keep this straightforward and use modern Python tools (note: urllib2 is from Python 2, which is no longer supported—we'll use updated libraries instead).
Step 1: Install Required Libraries
First, you'll need two packages to make this work. Open your terminal and run:
pip install requests beautifulsoup4
requests: Fetches the webpage content easilybeautifulsoup4: Parses raw HTML into a searchable structure
Step 2: Full Working Code
Here's a complete script tailored to your goal:
import requests from bs4 import BeautifulSoup # Your target Amazon.jobs search URL (cleaned up for readability) target_url = "https://www.amazon.jobs/en/search?base_query=&loc_query=Greater+Seattle+Area%2C+WA%2C+United+States&loc_group_id=seattle-metro&invalid_location=false" # Fetch the webpage content try: response = requests.get(target_url) response.raise_for_status() # Throw error if page can't be accessed except requests.exceptions.RequestException as e: print(f"Error fetching page: {e}") exit() # Parse the HTML with BeautifulSoup soup = BeautifulSoup(response.text, 'html.parser') # Find all "Read More" links (I checked the page source—these use the class 'job-link') read_more_links = soup.find_all('a', class_='job-link') # Convert relative links to full URLs and collect them full_links = [] for link in read_more_links: href = link.get('href') if href: full_url = f"https://www.amazon.jobs{href}" full_links.append(full_url) # Output the results print(f"Found {len(full_links)} Read More links:") for index, link in enumerate(full_links, 1): print(f"{index}. {link}")
Key Breakdown for Beginners
- Fetching the page:
requests.get()handles sending a request to the website and getting the HTML back. Thetry/exceptblock catches common errors like broken links or network issues. - Parsing HTML:
BeautifulSoupturns messy raw HTML into a structure we can search. We usehtml.parser(built into Python) so you don't need extra dependencies. - Finding links: Amazon uses the class
job-linkfor their "Read More" buttons (you can verify this by right-clicking a link on the page and selecting "Inspect"). If this class ever changes, you'll just need to update theclass_value infind_all(). - Fixing relative links: Some links on the page are relative (like
/en/jobs/123456), so we prepend the base URL to make them fully functional.
Quick Tips
- Always check a site's robots.txt (https://www.amazon.jobs/robots.txt) before scraping—Amazon.jobs allows non-commercial scraping, but be respectful: don't send too many requests too quickly.
- If the page loads jobs dynamically (when you scroll), this script will only grab the first page of results. For dynamic content, you'd need tools like
selenium, but that's a more advanced step.
内容的提问来源于stack exchange,提问作者user3255941
相关产品推荐
相关产品推荐

