使用Python检测Excel中URL的Active/Inactive状态并生成结果文件的技术求助
Hey there! Let's break down how to connect your URL status checker with Excel operations step by step—super doable even for a Python newbie 😊
First, we'll need two key tools to make this work smoothly:
pandas: Handles Excel file reading/writing like a prorequests: Checks if URLs are active (you can swap this with your existing working checker if you have one)openpyxl: Required for pandas to interact with .xlsx files
Install them via your terminal:
pip install pandas requests openpyxl
Assuming your input Excel has a single column named URL (if it's called something else, just adjust the column name in the code), here's how to read it into Python:
import pandas as pd # Load the input Excel file df = pd.read_excel("your_input_file.xlsx") # Optional: Verify the data loaded correctly (prints first 5 rows) print(df.head())
Let's create a function to check URL status (replace this with your existing working code if you have one!). This function returns "Active" if the URL loads successfully, or "Inactive" if it fails:
import requests def check_url_status(url): try: # Send a GET request with a timeout to avoid hanging response = requests.get(url, timeout=5) # Consider URLs with 2xx/3xx status codes as active if response.status_code in range(200, 400): return "Active" else: return "Inactive" except (requests.exceptions.ConnectionError, requests.exceptions.Timeout): # Handle cases where the URL can't be reached return "Inactive"
Now apply this function to every URL in your Excel column to create a new "Status" column:
# Add a new column with status results for each URL df["Status"] = df["URL"].apply(check_url_status)
Finally, save the updated data (with both URL and Status columns) to a new Excel file—this keeps your original file untouched:
# Save to a new Excel file (index=False removes pandas' default row numbers) df.to_excel("url_status_results.xlsx", index=False) print("Results saved successfully!")
- If your input Excel's URL column has a different name (e.g., "Links"), replace
df["URL"]withdf["Links"]in all code snippets. - If you have a huge list of URLs, the basic code might take time—for faster processing, you could explore parallel processing later, but start with this simple version first.
- Always test with a small sample of URLs first to make sure everything works as expected!
内容的提问来源于stack exchange,提问作者Lal Prashanth paulraj

