如何用Python3从Ethernodes提取表格至TXT?排查失效代码问题
解决Ethernodes节点数据提取问题
问题描述
需要从https://www.ethernodes.org/nodes提取节点数据并导出为TXT文件,方便bash脚本读取。但原有Python代码无法获取任何数据,代码如下:
import requests from bs4 import BeautifulSoup url = 'https://www.ethernodes.org/nodes?page=8' response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') host_ips = [] node_list = soup.find('ul', class_='nodes-list') if node_list is not None: for li in node_list.find_all('li'): host_ip = li.find('div', class_='node-host').text.strip() host_ips.append(host_ip) print(host_ips)
问题分析
- 页面结构变更:当前网站的节点列表不再使用
ul.nodes-list和div.node-host的HTML结构,原代码的选择器完全失效。 - 反爬拦截:网站会拦截无浏览器标识的请求,直接用
requests获取的内容不包含完整节点数据。 - 动态渲染:部分节点信息通过JavaScript动态加载,静态HTML无法覆盖全部内容。
修复方案
方法1:适配新HTML结构(静态可提取部分)
通过检查当前页面DOM,节点IP位于表格的指定列中,调整选择器并添加请求头即可提取:
import requests from bs4 import BeautifulSoup # 添加浏览器请求头,模拟正常访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } url = 'https://www.ethernodes.org/nodes?page=8' response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') host_ips = [] # 定位节点数据表格 node_table = soup.find('table', class_='table table-hover') if node_table: # 跳过表头,遍历数据行 for row in node_table.find_all('tr')[1:]: # 提取第二列的IP信息(根据页面实际列位置调整索引) ip_td = row.find_all('td')[1] if ip_td: host_ip = ip_td.text.strip() host_ips.append(host_ip) # 导出数据到TXT文件 with open('ethernodes_ips.txt', 'w') as f: for ip in host_ips: f.write(ip + '\n') print(f"已导出{len(host_ips)}个节点IP到ethernodes_ips.txt")
方法2:处理动态加载内容(完整数据)
如果静态HTML仍无法获取全部节点,使用selenium模拟浏览器加载动态内容:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup # 配置无头浏览器模式 chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome(options=chrome_options) url = 'https://www.ethernodes.org/nodes?page=8' driver.get(url) # 等待页面动态加载完成 driver.implicitly_wait(10) soup = BeautifulSoup(driver.page_source, 'html.parser') host_ips = [] node_table = soup.find('table', class_='table table-hover') if node_table: for row in node_table.find_all('tr')[1:]: ip_td = row.find_all('td')[1] if ip_td: host_ip = ip_td.text.strip() host_ips.append(host_ip) # 导出到TXT文件 with open('ethernodes_ips.txt', 'w') as f: for ip in host_ips: f.write(ip + '\n') driver.quit() print(f"已导出{len(host_ips)}个节点IP到ethernodes_ips.txt")
Bash脚本访问TXT文件
导出完成后,bash脚本可直接读取文件内容执行操作,示例:
# 遍历所有节点IP并执行自定义操作 while read ip; do echo "正在处理节点IP: $ip" # 此处添加你的业务逻辑,例如ping检测、端口扫描等 done < ethernodes_ips.txt
内容的提问来源于stack exchange,提问作者Magnus
相关产品推荐
相关产品推荐

