You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Apache Beam的read_utf8()触发TypeError的问题求助

解决Apache Beam读取GCS上gzip压缩文件的类型错误

问题出在你的文件是gzip压缩格式,但默认的ReadMatches没有启用自动解压逻辑,导致读取压缩文件流时内部状态异常,触发了TypeError: '<' not supported between instances of 'int' and 'NoneType'错误。

两种可行的解决方法

方法一:给ReadMatches指定压缩类型

修改ReadMatches步骤,添加compression_type参数明确指定GZIP压缩,让Beam自动处理解压:

import apache_beam as beam
from apache_beam.io import CompressionTypes
from apache_beam.io import fileio

with beam.Pipeline() as pipeline:
    readable_files = (
        pipeline
        | fileio.MatchFiles('<*filname.patterns>')
        # 指定压缩类型为GZIP
        | fileio.ReadMatches(compression_type=CompressionTypes.GZIP)
        | beam.Reshuffle())
    files_and_contents = (
        readable_files
        | beam.Map(lambda x: (x.metadata.path, x.read_utf8())))

方法二:使用ReadFromText(更适合文本场景)

如果你的文件是文本类gzip文件,推荐用ReadFromText,它会自动识别或处理压缩,代码更简洁:

import apache_beam as beam
from apache_beam.io import ReadFromText

with beam.Pipeline() as pipeline:
    files_and_contents = (
        pipeline
        # 指定压缩类型为gzip
        | ReadFromText('<*filname.patterns>', compression_type='gzip')
        # 带上文件路径元数据
        | beam.Map(lambda line, path: (path, line), with_metadata=True)
        # 按文件路径聚合所有行内容
        | beam.GroupByKey()
        | beam.Map(lambda item: (item[0], '\n'.join(item[1]))))

补充说明

gzip压缩文件的字节流结构和普通文本不同,默认读取逻辑未做解压处理时,会导致内部缓冲区的状态变量异常(出现None值),进而触发类型比较错误。只要明确指定压缩类型,Beam就会调用对应的解压逻辑,正常读取文件内容。

内容的提问来源于stack exchange,提问作者Daniel Abhishek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:40:30