Rails使用htmlentities gem时出现UTF-8无效字节序列问题
解决抓取网页时的"invalid byte sequence in UTF-8"错误
这个问题的核心不是数据库字段类型(你改LONGTEXT没用就验证了这点),而是你抓取到的网页内容编码和你预期的UTF-8不匹配,导致后续的htmlentities编码或字符串操作失败。像Google这类网站可能返回非UTF-8编码的内容,或者内容里夹杂了无效的UTF-8字节序列,直接处理就会触发错误。
下面是分步解决方案:
1. 正确识别并转换网页编码
抓取网页时,必须先确定目标页面的实际编码,再将其转换为标准UTF-8后再处理。这里以Ruby的Net::HTTP为例:
require 'net/http' require 'htmlentities' require 'uri' def fetch_and_encode_html(url_str) url = URI(url_str) response = Net::HTTP.get_response(url) # 从响应头或页面meta标签提取编码 encoding = if response['content-type']&.include?('charset') response['content-type'].match(/charset=([^;]+)/)[1] else # 如果头里没有,尝试解析页面meta标签(这里简化处理,实际可以用Nokogiri更准确) response.body.match(/<meta.*?charset=["']?([^"'>]+)["']?/i)&.[](1) || 'UTF-8' end # 强制按识别到的编码解析,再转成UTF-8,同时处理无效字符 begin utf8_content = response.body.force_encoding(encoding).encode( 'UTF-8', invalid: :replace, # 替换无效字节 undef: :replace, # 替换无法转换的字符 replace: '?' # 替换为问号,也可以用其他占位符 ) rescue Encoding::UndefinedConversionError, Encoding::InvalidByteSequenceError # 如果转码失败,直接清理无效UTF-8字节 utf8_content = response.body.scrub('?').encode('UTF-8') end # 用htmlentities编码 HTMLEntities.new.encode(utf8_content) end # 使用示例 encoded_html = fetch_and_encode_html('https://www.google.co.uk/') # 之后存入数据库...
2. 确保数据库连接使用完整UTF-8编码
即使字段是LONGTEXT,也要保证MySQL连接用的是utf8mb4(MySQL的utf8只支持部分Unicode字符,utf8mb4才是完整的UTF-8)。修改你的database.yml:
production: adapter: mysql2 database: your_database_name username: your_username password: your_password host: your_host encoding: utf8mb4 collation: utf8mb4_unicode_ci
3. 补充:用Nokogiri简化编码处理
如果用Nokogiri抓取网页,它会自动处理编码识别和转换,更省心:
require 'nokogiri' require 'open-uri' require 'htmlentities' doc = Nokogiri::HTML(URI.open('https://www.google.co.uk/')) # Nokogiri已经把内容转成UTF-8了 utf8_content = doc.to_html encoded_html = HTMLEntities.new.encode(utf8_content)
为什么lookagain.co.uk能正常运行?
因为这个网站的网页本身就是UTF-8编码,你的代码直接处理刚好匹配;而Google这类网站可能返回其他编码(比如ISO-8859-1)或者内容里有特殊字节,导致直接按UTF-8解析时出错。
内容的提问来源于stack exchange,提问作者Padu
相关产品推荐
相关产品推荐

