Flume Flow正则表达式报错:Must supply a valid regex string求助
解决Flume正则提取器的IllegalArgumentException错误
错误日志
java.lang.IllegalArgumentException: Must supply a valid regex string at com.google.common.base.Preconditions.checkArgument(Preconditions.java:141) at org.apache.flume.interceptor.RegexExtractorInterceptor$Builder.configure(RegexExtractorInterceptor.java:176) at org.apache.flume.channel.ChannelProcessor.configureInterceptors(ChannelProcessor.java:112) at org.apache.flume.channel.ChannelProcessor.configure(ChannelProcessor.java:82) at org.apache.flume.conf.Configurables.configure(Configurables.java:41) at org.apache.flume.node.AbstractConfigurationProvider.loadSources(AbstractConfigurationProvider.java:342) at org.apache.flume.node.AbstractConfigurationProvider.getConfiguration(AbstractConfigurationProvider.java:105) at org.apache.flume.node.PollingPropertiesFileConfigurationProvider$FileWatcherRunnable.run(PollingPropertiesFileConfigurationProvider.java:145) at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) at java.lang.Thread.run(Thread.java:750)
原始Flume配置文件
# Purpose: load data to hdfs and partition it by year and month # Name source, sink and channel; replicate to logger so we can see the timestamp header agent1.sources = source2 agent1.sinks = sink-hdfs agent1.channels = channel-hdfs # Describe and configure the source. agent1.sources.source2.type = spooldir agent1.sources.source2.spoolDir = /home/hadoopuser/Task4/spooldir agent1.sources.source2.interceptors = i2 i3 i4 agent1.sources.source2.interceptors.i2.type = regex_extractor agent1.sources.source2.interceptors.i3.type = regex_extractor agent1.sources.source2.interceptors.i4.type = regex_extractor # regex to pick up the year agent1.sources.source2.interceptors.i2.regex = (?<=\\s)[0-9]{4}(?=-) agent1.sources.source2.interceptors.i2.serializers = y agent1.sources.source2.interceptors.i2.serializers.y.name = year #regex for year #agent1.sources.source2.interceptors.ye.regex = ^[0-9]{3}[a-zA-Z0-9] #agent1.sources.source2.interceptors.ye.serializers = y1 #agent1.sources.source2.interceptors.ye.serializers.y1.name = year # regex to pick up the month agent1.sources.source2.interceptors.i3.regex = (?<=-)[0-9]{2}(?=-) agent1.sources.source2.interceptors.i3.serializers = m agent1.sources.source2.interceptors.i3.serializers.m.name = month #regex for mointhy #agent1.sources.source2.interceptors.mo.regex = [0-9]+ #agent1.sources.source2.interceptors.mo.serializers = m1 #agent1.sources.source2.interceptors.mo.serializers.m1.name = month # Define the HDFS sink 2 –year and month agent1.sinks.sink-hdfs.type = hdfs agent1.sinks.sink-hdfs.hdfs.path = /Task4/partA_flume/%{year}/%{month} agent1.sinks.sink-hdfs.hdfs.filePrefix = %{year}-%{month} agent1.sinks.sink-hdfs.hdfs.fileSuffix = .txt # Bind the source and sinks to the channels agent1.sources.source2.channels = channel-hdfs agent1.sinks.sink-hdfs.channel = channel-hdfs # The channel will buffer events to file for durability. Type memory is faster but volatile. agent1.channels.channel-hdfs.type = memory # -- end of file
待加载数据格式
5016833 1 2014-01-02 15:38:40 20719.257632 0 5016834 1 2014-01-02 15:38:50 20719.262176 0 5016835 1 2014-01-02 15:39:00 20719.26672 0 5016836 1 2014-01-02 15:39:10 20719.271264 0
问题原因与修复方案
直接报错原因
配置中声明了i4这个regex_extractor类型的拦截器,但未给它配置regex参数,Flume初始化时会检查每个正则提取器的正则表达式是否合法,缺失参数导致抛出IllegalArgumentException。
修复步骤
移除未配置的i4拦截器
将配置中的拦截器列表修改为:agent1.sources.source2.interceptors = i2 i3同时删除无用的
agent1.sources.source2.interceptors.i4.type = regex_extractor配置行。优化正则提取逻辑(可选但推荐)
原配置用两个拦截器分别提取年和月,可改为用一个拦截器一次性提取,减少配置复杂度和运行开销:agent1.sources.source2.interceptors = ym agent1.sources.source2.interceptors.ym.type = regex_extractor # 匹配数据中的yyyy-MM-dd部分,捕获年和月两个分组 agent1.sources.source2.interceptors.ym.regex = \\s(\\d{4})-(\\d{2})-\\d{2}\\s agent1.sources.source2.interceptors.ym.serializers = y m agent1.sources.source2.interceptors.ym.serializers.y.name = year agent1.sources.source2.interceptors.ym.serializers.m.name = month这个正则会匹配数据中日期字段的空白符、4位年份、2位月份,通过分组直接提取年和月,无需多个拦截器。
验证正则转义
原配置中的正则转义(如(?<=\\s))在properties文件中是正确的,因为properties文件中\需要用\\转义,确保Java正则引擎能正确识别空白符等特殊字符。
修改后的完整配置
# Purpose: load data to hdfs and partition it by year and month agent1.sources = source2 agent1.sinks = sink-hdfs agent1.channels = channel-hdfs # Source configuration agent1.sources.source2.type = spooldir agent1.sources.source2.spoolDir = /home/hadoopuser/Task4/spooldir # Interceptor: extract year and month in one step agent1.sources.source2.interceptors = ym agent1.sources.source2.interceptors.ym.type = regex_extractor agent1.sources.source2.interceptors.ym.regex = \\s(\\d{4})-(\\d{2})-\\d{2}\\s agent1.sources.source2.interceptors.ym.serializers = y m agent1.sources.source2.interceptors.ym.serializers.y.name = year agent1.sources.source2.interceptors.ym.serializers.m.name = month # HDFS Sink configuration agent1.sinks.sink-hdfs.type = hdfs agent1.sinks.sink-hdfs.hdfs.path = /Task4/partA_flume/%{year}/%{month} agent1.sinks.sink-hdfs.hdfs.filePrefix = %{year}-%{month} agent1.sinks.sink-hdfs.hdfs.fileSuffix = .txt # Channel binding agent1.sources.source2.channels = channel-hdfs agent1.sinks.sink-hdfs.channel = channel-hdfs # Channel configuration agent1.channels.channel-hdfs.type = memory
内容的提问来源于stack exchange,提问作者coder12345
相关产品推荐
相关产品推荐

