You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Nokogiri(Ruby)从多元素列提取单个加密货币名称

解决网页抓取中提取加密货币全名的问题

Hey there! Let's fix this up for you step by step. First, I noticed a tiny typo in your code: coin_table.cc("tr") should be coin_table.css("tr") since cc isn't a valid Nokogiri method—easy mistake when you're starting out!

Now, onto extracting just the full cryptocurrency names. Your current output has extra newlines, spaces, and both the ticker (like BTC) and full name (like Bitcoin). Here are two solid approaches to get what you need:

方法1:清理文本并拆分(适配当前输出情况)

We'll first strip all extra whitespace, then split the text to isolate the full name:

site = "http://www.cyrptomarketcap.com"
doc = Nokogiri::HTML.parse(open(site))
coin_table = doc.css("table").sort { |x,y| y.css("tr").count <=> x.css("tr").count }.first
rows = coin_table.css("tr") # 修正了这里的拼写错误
rows = rows.select { |row| row.css("th").empty? }

data = rows.map do |row|
  cell_text = row.at_css("td:nth-child(2)").try(:text)
  next [] unless cell_text # 跳过没有文本的单元格
  
  # 去除前后的空白和换行符
  cleaned = cell_text.strip
  # 按换行拆分文本,取最后一段(也就是全名),再做一次空白清理
  coin_name = cleaned.split("\n").last.strip
  [coin_name]
end

运行后你会得到类似 [["Bitcoin"], ["Ripple"], ...] 的结果——正好是你想要的加密货币全名。

方法2:直接定位HTML子元素(更可靠)

如果你检查那个<td>元素的HTML结构,大概率会发现代码缩写和全名分别放在不同的子标签里(比如<span>)。举个例子,如果全名在带有coin-name类的标签里,你可以直接定位这个元素:

# 把数据映射部分替换成这段代码
data = rows.map do |row|
  coin_name = row.at_css("td:nth-child(2) .coin-name").try(:text)
  [coin_name&.strip].compact # 处理空值并清理空白
end

这种方法更稳健,因为它不依赖文本格式,而是直接抓取存储全名的精确元素。

小提示:用浏览器的开发者工具(右键→检查)查看表格单元格的HTML结构,这能帮你找到最精准的选择器。

内容的提问来源于stack exchange,提问作者mattC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:26:18