Python代码修改:搜索网页源码、读取URL列表并匹配关键词
Solution: Modified Code to Search Web Source Code & Read URLs from File
Let's tweak your code to hit all your requirements—searching the raw web source code, ditching the count, reading URLs from a text file, and outputting a clear message when the keyword is found. Here's the revised version with explanations:
import requests def check_keyword_in_source(url, keyword): # Add http/https prefix if missing to avoid requests errors if not url.startswith(('http://', 'https://')): url = f'https://{url}' try: # Fetch the raw web source code directly response = requests.get(url, allow_redirects=False, timeout=10) response.raise_for_status() # Catch HTTP errors like 404 or 500 # Check if keyword exists in the unprocessed source code if keyword in response.text: print(f"{keyword} Found on this website") except requests.exceptions.RequestException as e: print(f"Failed to access {url}: {str(e)}") def main(): # Replace with your text file path (one URL per line) urls_file = 'urls.txt' target_keyword = 'google' # Update this to your desired keyword # Read and process URLs from the file try: with open(urls_file, 'r') as file: # Skip empty lines and clean up whitespace from each URL urls = [line.strip() for line in file if line.strip()] for url in urls: check_keyword_in_source(url, target_keyword) except FileNotFoundError: print(f"Error: Could not find the file {urls_file}") if __name__ == '__main__': main()
Key Updates Explained:
- Removed BeautifulSoup: We don't need HTML parsing anymore—
response.textgives us the full, unprocessed web source code directly, which is exactly what you want to search. - URL Fixer: Added a quick check to prepend
https://to URLs without an HTTP/HTTPS prefix (prevents invalid URL errors in requests). - File Reading: The script pulls URLs from a text file (one per line), skips empty lines, and processes each entry automatically.
- Error Handling: Added
try-exceptblocks to handle common issues like missing files, failed connections, or broken URLs—so the script won't crash if one URL has problems. - Counting Removed: Replaced the count functionality with a simple existence check that prints your requested message when the keyword is found.
How to Use:
- Create a text file (e.g.,
urls.txt) and add one URL per line (likegoogle.com,example.com). - Update
target_keywordin themain()function to your search term. - Run the script—it will loop through every URL, check the source code, and print results as it goes.
内容的提问来源于stack exchange,提问作者JDoe
相关产品推荐
相关产品推荐

