You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB中按多字符分割字符串并通过MapReduce统计同长度词频

Step-by-Step Implementation for Word Length Count with MapReduce

Let’s break down how to solve this problem—from cleaning your input string to implementing the MapReduce job that counts words by their length, matching your expected result.

1. First: Clean the Input String

Your sentence has punctuation (--, ?, -) that we need to strip to get valid words. The key here is to extract only alphabetic sequences (since words are made of letters). Using a regex like [a-zA-Z]+ will pull out all valid words and ignore punctuation/empty strings from splits.

For your input:

SUPPOSING that Truth is a woman--what then? I will get -

This regex gives us:
SUPPOSING, that, Truth, is, a, woman, what, then, I, will, get

(Note: Your expected result omits I, will, get—likely a typo in the example, but our solution will handle all valid words correctly.)

2. MapReduce Implementation

Mapper Class (Java Example)

The mapper takes each line of input, splits it into clean words, calculates each word’s length, and emits pairs like (length, 1) (each word contributes 1 to its length’s count).

import java.io.IOException;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Mapper;

public class WordLengthCountMapper extends Mapper<LongWritable, Text, IntWritable, IntWritable> {
    private final static IntWritable one = new IntWritable(1);
    private IntWritable wordLength = new IntWritable();

    @Override
    protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException {
        // Split line into valid words using regex
        String[] words = value.toString().split("[^a-zA-Z]+");
        
        for (String word : words) {
            if (!word.isEmpty()) { // Skip empty strings from punctuation splits
                wordLength.set(word.length());
                context.write(wordLength, one);
            }
        }
    }
}

Reducer Class (Java Example)

The reducer aggregates all the 1 values for each length and sums them up, emitting the final count per length.

import java.io.IOException;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.mapreduce.Reducer;

public class WordLengthCountReducer extends Reducer<IntWritable, IntWritable, IntWritable, IntWritable> {
    @Override
    protected void reduce(IntWritable key, Iterable<IntWritable> values, Context context) throws IOException, InterruptedException {
        int totalCount = 0;
        for (IntWritable count : values) {
            totalCount += count.get();
        }
        context.write(key, new IntWritable(totalCount));
    }
}

Driver Class (Java Example)

This sets up the MapReduce job, linking the mapper/reducer, defining input/output paths, and configuring data types.

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

public class WordLengthCountDriver {
    public static void main(String[] args) throws Exception {
        Configuration conf = new Configuration();
        Job job = Job.getInstance(conf, "word-length-count");
        
        job.setJarByClass(WordLengthCountDriver.class);
        job.setMapperClass(WordLengthCountMapper.class);
        job.setCombinerClass(WordLengthCountReducer.class); // Optional: Sums counts at mapper to reduce data transfer
        job.setReducerClass(WordLengthCountReducer.class);
        
        job.setOutputKeyClass(IntWritable.class);
        job.setOutputValueClass(IntWritable.class);
        
        FileInputFormat.addInputPath(job, new Path(args[0]));
        FileOutputFormat.setOutputPath(job, new Path(args[1]));
        
        System.exit(job.waitForCompletion(true) ? 0 : 1);
    }
}

3. Python Alternative (Hadoop Streaming)

If you prefer Python, here’s how to implement the same logic with Hadoop Streaming:

Mapper Script

import sys
import re

for line in sys.stdin:
    line = line.strip()
    # Extract all alphabetic words
    words = re.findall(r'[a-zA-Z]+', line)
    for word in words:
        print(f"{len(word)}\t1")

Reducer Script

import sys

current_length = None
current_count = 0

for line in sys.stdin:
    line = line.strip()
    length, count = line.split('\t', 1)
    try:
        count = int(count)
        length = int(length)
    except ValueError:
        continue  # Skip invalid lines
    
    if current_length == length:
        current_count += count
    else:
        if current_length is not None:
            print(f"{current_length}\t{current_count}")
        current_length = length
        current_count = count

# Print the last length count
if current_length is not None:
    print(f"{current_length}\t{current_count}")

4. Get Your Expected JSON Output

After running the MapReduce job, you’ll get a text file with lines like:

1	1
2	1
4	3
5	2
9	1

You can easily convert this into your desired JSON format with a simple script (e.g., in Python) that reads the output and constructs the array of objects.

内容的提问来源于stack exchange,提问作者BarZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:10:45