You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ruby维基百科爬取程序优化:实现流式输入触发爬取

优化Ruby维基百科爬虫的流式输入自动触发逻辑

Hey there! Sounds like you want to make your Wikipedia scraper feel way more responsive—letting users type naturally, and auto-triggering the search when they pause for a second instead of forcing them to input in chunks. Let's break down how to pull this off smoothly in Ruby.

核心思路

We need two key pieces to make this work:

  • Real-time character capture: Grab each character as the user types (no waiting for Enter presses)
  • Idle timeout detection: Reset a timer every time the user types, and fire off the search when the timer hits 1 second of inactivity

完整实现代码

Here's a working example with comments walking through the important bits:

require 'io/console'
require 'net/http'
require 'json'

# 存储用户输入的缓冲区
@input_buffer = ""
# 记录最后一次输入的时间戳
@last_input_time = Time.now
# 控制后台线程的运行状态
@running = true

# 后台线程:检测输入停顿,触发爬取逻辑
Thread.new do
  while @running
    current_time = Time.now
    # 当停顿超过1秒且输入不为空时触发爬取
    if (current_time - @last_input_time) >= 1 && !@input_buffer.strip.empty?
      puts "\n--- 正在爬取「#{@input_buffer.strip}」的内容 ---"
      
      # 调用维基百科搜索API(URL编码避免特殊字符问题)
      search_query = URI.encode_www_form_component(@input_buffer.strip)
      url = URI("https://en.wikipedia.org/w/api.php?action=query&list=search&srsearch=#{search_query}&format=json")
      response = Net::HTTP.get(url)
      data = JSON.parse(response)
      
      # 解析并格式化输出结果
      if data.dig('query', 'search').any?
        puts "搜索结果:"
        data['query']['search'].each_with_index do |item, idx|
          # 移除HTML标签,显示纯文本摘要
          clean_snippet = item['snippet'].gsub(/<[^>]+>/, '')
          puts "#{idx+1}. #{item['title']} - #{clean_snippet}"
        end
      else
        puts "未找到相关内容"
      end
      puts "\n继续输入(按Ctrl+C退出):"
      # 清空缓冲区,准备下一轮输入
      @input_buffer = ""
    end
    sleep 0.1 # 降低后台线程的CPU占用
  end
end

# 主线程:处理实时输入
puts "开始输入内容,停顿1秒自动触发爬取(按Ctrl+C退出):"
begin
  while @running
    char = STDIN.getch
    case char
    when "\b" # 处理退格键,优化输入体验
      unless @input_buffer.empty?
        @input_buffer.chop!
        # 在控制台回退并删除字符
        print "\b \b"
      end
    when "\r", "\n" # 忽略回车换行符
      next
    else
      @input_buffer << char
      print char # 在控制台实时显示输入的字符
      @last_input_time = Time.now # 重置最后输入时间
    end
  end
rescue Interrupt
  @running = false
  puts "\n程序已退出"
end

关键细节说明

  • 实时输入处理: We use STDIN.getch from the io/console module to read each character the moment the user presses it—no more waiting for them to hit Enter.
  • Idle timer: The background thread checks every 0.1 seconds if the user has stopped typing for 1 full second. If so, it kicks off the Wikipedia API call automatically.
  • User experience tweaks: We added backspace support so users can correct typos, and clear the input buffer after each search to let them start fresh immediately.
  • API safety: We URL-encode the search query to avoid errors with special characters like spaces or punctuation.

This setup lets users type naturally, and the scraper will run the search as soon as they pause—no more tedious chunked input waits!

内容的提问来源于stack exchange,提问作者Rishav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:16:33