You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python3从Ethernodes提取表格至TXT?排查失效代码问题

解决Ethernodes节点数据提取问题

问题描述

需要从https://www.ethernodes.org/nodes提取节点数据并导出为TXT文件,方便bash脚本读取。但原有Python代码无法获取任何数据,代码如下:

import requests
from bs4 import BeautifulSoup

url = 'https://www.ethernodes.org/nodes?page=8'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

host_ips = []
node_list = soup.find('ul', class_='nodes-list')
if node_list is not None:
    for li in node_list.find_all('li'):
        host_ip = li.find('div', class_='node-host').text.strip()
        host_ips.append(host_ip)

print(host_ips)

问题分析

  1. 页面结构变更:当前网站的节点列表不再使用ul.nodes-list和div.node-host的HTML结构,原代码的选择器完全失效。
  2. 反爬拦截:网站会拦截无浏览器标识的请求,直接用requests获取的内容不包含完整节点数据。
  3. 动态渲染:部分节点信息通过JavaScript动态加载,静态HTML无法覆盖全部内容。

修复方案

方法1:适配新HTML结构(静态可提取部分)

通过检查当前页面DOM,节点IP位于表格的指定列中,调整选择器并添加请求头即可提取:

import requests
from bs4 import BeautifulSoup

# 添加浏览器请求头,模拟正常访问
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

url = 'https://www.ethernodes.org/nodes?page=8'
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

host_ips = []
# 定位节点数据表格
node_table = soup.find('table', class_='table table-hover')
if node_table:
    # 跳过表头,遍历数据行
    for row in node_table.find_all('tr')[1:]:
        # 提取第二列的IP信息(根据页面实际列位置调整索引)
        ip_td = row.find_all('td')[1]
        if ip_td:
            host_ip = ip_td.text.strip()
            host_ips.append(host_ip)

# 导出数据到TXT文件
with open('ethernodes_ips.txt', 'w') as f:
    for ip in host_ips:
        f.write(ip + '\n')

print(f"已导出{len(host_ips)}个节点IP到ethernodes_ips.txt")

方法2:处理动态加载内容(完整数据)

如果静态HTML仍无法获取全部节点,使用selenium模拟浏览器加载动态内容:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

# 配置无头浏览器模式
chrome_options = Options()
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument('--disable-dev-shm-usage')

driver = webdriver.Chrome(options=chrome_options)
url = 'https://www.ethernodes.org/nodes?page=8'
driver.get(url)

# 等待页面动态加载完成
driver.implicitly_wait(10)

soup = BeautifulSoup(driver.page_source, 'html.parser')
host_ips = []

node_table = soup.find('table', class_='table table-hover')
if node_table:
    for row in node_table.find_all('tr')[1:]:
        ip_td = row.find_all('td')[1]
        if ip_td:
            host_ip = ip_td.text.strip()
            host_ips.append(host_ip)

# 导出到TXT文件
with open('ethernodes_ips.txt', 'w') as f:
    for ip in host_ips:
        f.write(ip + '\n')

driver.quit()
print(f"已导出{len(host_ips)}个节点IP到ethernodes_ips.txt")

Bash脚本访问TXT文件

导出完成后,bash脚本可直接读取文件内容执行操作,示例:

# 遍历所有节点IP并执行自定义操作
while read ip; do
    echo "正在处理节点IP: $ip"
    # 此处添加你的业务逻辑,例如ping检测、端口扫描等
done < ethernodes_ips.txt

内容的提问来源于stack exchange,提问作者Magnus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 10:27:59