You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Rails使用htmlentities gem时出现UTF-8无效字节序列问题

解决抓取网页时的"invalid byte sequence in UTF-8"错误

这个问题的核心不是数据库字段类型(你改LONGTEXT没用就验证了这点),而是你抓取到的网页内容编码和你预期的UTF-8不匹配,导致后续的htmlentities编码或字符串操作失败。像Google这类网站可能返回非UTF-8编码的内容,或者内容里夹杂了无效的UTF-8字节序列,直接处理就会触发错误。

下面是分步解决方案:

1. 正确识别并转换网页编码

抓取网页时,必须先确定目标页面的实际编码,再将其转换为标准UTF-8后再处理。这里以Ruby的Net::HTTP为例:

require 'net/http'
require 'htmlentities'
require 'uri'

def fetch_and_encode_html(url_str)
  url = URI(url_str)
  response = Net::HTTP.get_response(url)

  # 从响应头或页面meta标签提取编码
  encoding = if response['content-type']&.include?('charset')
               response['content-type'].match(/charset=([^;]+)/)[1]
             else
               # 如果头里没有,尝试解析页面meta标签(这里简化处理,实际可以用Nokogiri更准确)
               response.body.match(/<meta.*?charset=["']?([^"'>]+)["']?/i)&.[](1) || 'UTF-8'
             end

  # 强制按识别到的编码解析,再转成UTF-8,同时处理无效字符
  begin
    utf8_content = response.body.force_encoding(encoding).encode(
      'UTF-8',
      invalid: :replace,  # 替换无效字节
      undef: :replace,    # 替换无法转换的字符
      replace: '?'        # 替换为问号,也可以用其他占位符
    )
  rescue Encoding::UndefinedConversionError, Encoding::InvalidByteSequenceError
    # 如果转码失败,直接清理无效UTF-8字节
    utf8_content = response.body.scrub('?').encode('UTF-8')
  end

  # 用htmlentities编码
  HTMLEntities.new.encode(utf8_content)
end

# 使用示例
encoded_html = fetch_and_encode_html('https://www.google.co.uk/')
# 之后存入数据库...

2. 确保数据库连接使用完整UTF-8编码

即使字段是LONGTEXT,也要保证MySQL连接用的是utf8mb4(MySQL的utf8只支持部分Unicode字符,utf8mb4才是完整的UTF-8)。修改你的database.yml:

production:
  adapter: mysql2
  database: your_database_name
  username: your_username
  password: your_password
  host: your_host
  encoding: utf8mb4
  collation: utf8mb4_unicode_ci

3. 补充:用Nokogiri简化编码处理

如果用Nokogiri抓取网页,它会自动处理编码识别和转换,更省心:

require 'nokogiri'
require 'open-uri'
require 'htmlentities'

doc = Nokogiri::HTML(URI.open('https://www.google.co.uk/'))
# Nokogiri已经把内容转成UTF-8了
utf8_content = doc.to_html
encoded_html = HTMLEntities.new.encode(utf8_content)

为什么lookagain.co.uk能正常运行?

因为这个网站的网页本身就是UTF-8编码,你的代码直接处理刚好匹配;而Google这类网站可能返回其他编码(比如ISO-8859-1)或者内容里有特殊字节,导致直接按UTF-8解析时出错。

内容的提问来源于stack exchange,提问作者Padu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:48:37