使用Apache Beam的read_utf8()触发TypeError的问题求助
解决Apache Beam读取GCS上gzip压缩文件的类型错误
问题出在你的文件是gzip压缩格式,但默认的ReadMatches没有启用自动解压逻辑,导致读取压缩文件流时内部状态异常,触发了TypeError: '<' not supported between instances of 'int' and 'NoneType'错误。
两种可行的解决方法
方法一:给ReadMatches指定压缩类型
修改ReadMatches步骤,添加compression_type参数明确指定GZIP压缩,让Beam自动处理解压:
import apache_beam as beam from apache_beam.io import CompressionTypes from apache_beam.io import fileio with beam.Pipeline() as pipeline: readable_files = ( pipeline | fileio.MatchFiles('<*filname.patterns>') # 指定压缩类型为GZIP | fileio.ReadMatches(compression_type=CompressionTypes.GZIP) | beam.Reshuffle()) files_and_contents = ( readable_files | beam.Map(lambda x: (x.metadata.path, x.read_utf8())))
方法二:使用ReadFromText(更适合文本场景)
如果你的文件是文本类gzip文件,推荐用ReadFromText,它会自动识别或处理压缩,代码更简洁:
import apache_beam as beam from apache_beam.io import ReadFromText with beam.Pipeline() as pipeline: files_and_contents = ( pipeline # 指定压缩类型为gzip | ReadFromText('<*filname.patterns>', compression_type='gzip') # 带上文件路径元数据 | beam.Map(lambda line, path: (path, line), with_metadata=True) # 按文件路径聚合所有行内容 | beam.GroupByKey() | beam.Map(lambda item: (item[0], '\n'.join(item[1]))))
补充说明
gzip压缩文件的字节流结构和普通文本不同,默认读取逻辑未做解压处理时,会导致内部缓冲区的状态变量异常(出现None值),进而触发类型比较错误。只要明确指定压缩类型,Beam就会调用对应的解压逻辑,正常读取文件内容。
内容的提问来源于stack exchange,提问作者Daniel Abhishek
相关产品推荐
相关产品推荐

