You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将文本文件中的URL读取到Python的BeautifulSoup脚本中?

How to Modify the Scraper to Read URLs from a Text File

Got it, let's tweak your existing script to pull URLs from a newline-separated text file instead of hardcoding a single one. This approach will scale easily for your 9 target URLs, and we'll add a bit of error handling to keep things robust.

Modified Script

from bs4 import BeautifulSoup
import requests

# Define the path to your URL text file
url_file_path = "urls.txt"

# Read URLs from the file, clean up whitespace and skip empty lines
with open(url_file_path, 'r') as file:
    urls = [line.strip() for line in file if line.strip()]

# Loop through each URL and run the scraping logic
for url in urls:
    try:
        print(f"Scraping {url}...")
        response = requests.get(url)
        # Check if the request was successful (status code 200)
        response.raise_for_status()
        
        data = response.text
        soup = BeautifulSoup(data, 'html.parser')
        categories = soup.find_all("a", {"class":'navlabellink nvoffset nnormal'})
        
        # Print results for each category found
        for category in categories:
            print(f"{url},{category.text}")
            
    except requests.exceptions.RequestException as e:
        print(f"Failed to scrape {url}: {str(e)}")

Key Changes Explained

  • Reading the URL File: We use a with statement (best practice for file handling) to open the text file. The list comprehension [line.strip() for line in file if line.strip()] cleans up each line by removing newlines/whitespace and skips any empty lines in the file.
  • Loop Through URLs: The scraping logic that was tied to a single url variable is now wrapped in a loop over all URLs from the file.
  • Error Handling: The try-except block catches common request issues (like broken links, timeouts, or HTTP errors) so a single failed URL won't crash the entire script. We also use response.raise_for_status() to explicitly check if the HTTP request succeeded.
  • Clear Output: Added a print statement to show which URL is being scraped, making it easier to track progress with 9 URLs.

Notes for Your Setup

  • Make sure your urls.txt file is in the same directory as the script, or use the full file path (like C:/scraper/urls.txt on Windows or /home/user/scraper/urls.txt on Linux/macOS).
  • Each URL in the file should be on its own line, e.g.:
    http://www.url1.com
    http://www.url2.com
    http://www.url3.com
    ...
    

内容的提问来源于stack exchange,提问作者mjbaybay7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:17:31