Ubuntu下运行Python URL提取代码遇URLError及Anaconda启动缓慢问题
Hey there, let's break down and fix both problems you're facing: the URLError stopping your URL scraping, and the glacial Anaconda startup time.
1. Fixing URLError: [Errno 110] Connection Timed Out
Looking at your stack trace, the core issue is that your Python script can't establish a connection to the target URLs before timing out. Here's how to resolve this:
Clean up your URL inputs
When usingreadline()to read fromurl.txt, each line likely includes a trailing newline character (\n), which breaks the URL. Add.strip()to remove extra whitespace:line = file1.readline().strip() if not line: breakAdd a timeout to your requests
The defaulturlopencall doesn't have a timeout, so it can hang indefinitely. Set a reasonable timeout (e.g., 10 seconds) to avoid this:response = urllib.request.urlopen(line, context=gcontext, timeout=10)Handle proxies (if your network uses them)
If your Ubuntu system requires a proxy to access external URLs, configure it in your script:# Replace with your proxy details proxy_handler = urllib.request.ProxyHandler({ 'http': 'http://your-proxy-address:port', 'https': 'https://your-proxy-address:port' }) opener = urllib.request.build_opener(proxy_handler) urllib.request.install_opener(opener)Verify URL accessibility from your terminal
Test if the URL works outside Python usingcurlto rule out network/firewall issues:curl https://your-target-url.com # Replace with a URL from your url.txtTweak your SSL context (if needed)
Your currentSSLContext()might have compatibility issues with some sites. Try explicitly setting a TLS protocol:gcontext = ssl.SSLContext(ssl.PROTOCOL_TLS_CLIENT) # Only use the below two lines if you trust the target sites (not recommended for production) gcontext.check_hostname = False gcontext.verify_mode = ssl.CERT_NONE
2. Speeding Up Anaconda's 20-Minute Startup
Anaconda's slow startup is almost always related to automatic initialization or bloated cache. Try these fixes:
Disable automatic conda initialization
By default, conda adds code to your shell config (~/.bashrcor~/.zshrc) that runs every time you open a terminal. Comment out those lines to skip auto-loading:# Open your shell config file nano ~/.bashrcLook for and comment out the conda initialization block (it starts with
__conda_setup=...). After saving, you can manually activate conda when needed with:source /home/narendra/anaconda3/bin/activateClear conda's cache
Old package caches can slow down conda operations. Clean them up with:conda clean -aOptimize conda with faster mirrors
Using domestic mirrors can speed up package searches. Create or edit~/.condarcwith:channels: - defaults show_channel_urls: true default_channels: - https://mirrors.tuna.tsinghua.edu.cn/anaconda/pkgs/main - https://mirrors.tuna.tsinghua.edu.cn/anaconda/pkgs/r - https://mirrors.tuna.tsinghua.edu.cn/anaconda/pkgs/msys2 custom_channels: conda-forge: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud bioconda: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud pytorch: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloudCheck system resource usage
If your CPU or RAM is maxed out when starting Anaconda, close unnecessary background processes. Usehtopin the terminal to monitor resource usage during startup.
内容的提问来源于stack exchange,提问作者Nitesh

