如何将文本文件中的URL读取到Python的BeautifulSoup脚本中?
How to Modify the Scraper to Read URLs from a Text File
Got it, let's tweak your existing script to pull URLs from a newline-separated text file instead of hardcoding a single one. This approach will scale easily for your 9 target URLs, and we'll add a bit of error handling to keep things robust.
Modified Script
from bs4 import BeautifulSoup import requests # Define the path to your URL text file url_file_path = "urls.txt" # Read URLs from the file, clean up whitespace and skip empty lines with open(url_file_path, 'r') as file: urls = [line.strip() for line in file if line.strip()] # Loop through each URL and run the scraping logic for url in urls: try: print(f"Scraping {url}...") response = requests.get(url) # Check if the request was successful (status code 200) response.raise_for_status() data = response.text soup = BeautifulSoup(data, 'html.parser') categories = soup.find_all("a", {"class":'navlabellink nvoffset nnormal'}) # Print results for each category found for category in categories: print(f"{url},{category.text}") except requests.exceptions.RequestException as e: print(f"Failed to scrape {url}: {str(e)}")
Key Changes Explained
- Reading the URL File: We use a
withstatement (best practice for file handling) to open the text file. The list comprehension[line.strip() for line in file if line.strip()]cleans up each line by removing newlines/whitespace and skips any empty lines in the file. - Loop Through URLs: The scraping logic that was tied to a single
urlvariable is now wrapped in a loop over all URLs from the file. - Error Handling: The
try-exceptblock catches common request issues (like broken links, timeouts, or HTTP errors) so a single failed URL won't crash the entire script. We also useresponse.raise_for_status()to explicitly check if the HTTP request succeeded. - Clear Output: Added a print statement to show which URL is being scraped, making it easier to track progress with 9 URLs.
Notes for Your Setup
- Make sure your
urls.txtfile is in the same directory as the script, or use the full file path (likeC:/scraper/urls.txton Windows or/home/user/scraper/urls.txton Linux/macOS). - Each URL in the file should be on its own line, e.g.:
http://www.url1.com http://www.url2.com http://www.url3.com ...
内容的提问来源于stack exchange,提问作者mjbaybay7
相关产品推荐
相关产品推荐

