MongoDB中按多字符分割字符串并通过MapReduce统计同长度词频
Let’s break down how to solve this problem—from cleaning your input string to implementing the MapReduce job that counts words by their length, matching your expected result.
1. First: Clean the Input String
Your sentence has punctuation (--, ?, -) that we need to strip to get valid words. The key here is to extract only alphabetic sequences (since words are made of letters). Using a regex like [a-zA-Z]+ will pull out all valid words and ignore punctuation/empty strings from splits.
For your input:
SUPPOSING that Truth is a woman--what then? I will get -
This regex gives us:SUPPOSING, that, Truth, is, a, woman, what, then, I, will, get
(Note: Your expected result omits I, will, get—likely a typo in the example, but our solution will handle all valid words correctly.)
2. MapReduce Implementation
Mapper Class (Java Example)
The mapper takes each line of input, splits it into clean words, calculates each word’s length, and emits pairs like (length, 1) (each word contributes 1 to its length’s count).
import java.io.IOException; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.LongWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.mapreduce.Mapper; public class WordLengthCountMapper extends Mapper<LongWritable, Text, IntWritable, IntWritable> { private final static IntWritable one = new IntWritable(1); private IntWritable wordLength = new IntWritable(); @Override protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException { // Split line into valid words using regex String[] words = value.toString().split("[^a-zA-Z]+"); for (String word : words) { if (!word.isEmpty()) { // Skip empty strings from punctuation splits wordLength.set(word.length()); context.write(wordLength, one); } } } }
Reducer Class (Java Example)
The reducer aggregates all the 1 values for each length and sums them up, emitting the final count per length.
import java.io.IOException; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.mapreduce.Reducer; public class WordLengthCountReducer extends Reducer<IntWritable, IntWritable, IntWritable, IntWritable> { @Override protected void reduce(IntWritable key, Iterable<IntWritable> values, Context context) throws IOException, InterruptedException { int totalCount = 0; for (IntWritable count : values) { totalCount += count.get(); } context.write(key, new IntWritable(totalCount)); } }
Driver Class (Java Example)
This sets up the MapReduce job, linking the mapper/reducer, defining input/output paths, and configuring data types.
import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; public class WordLengthCountDriver { public static void main(String[] args) throws Exception { Configuration conf = new Configuration(); Job job = Job.getInstance(conf, "word-length-count"); job.setJarByClass(WordLengthCountDriver.class); job.setMapperClass(WordLengthCountMapper.class); job.setCombinerClass(WordLengthCountReducer.class); // Optional: Sums counts at mapper to reduce data transfer job.setReducerClass(WordLengthCountReducer.class); job.setOutputKeyClass(IntWritable.class); job.setOutputValueClass(IntWritable.class); FileInputFormat.addInputPath(job, new Path(args[0])); FileOutputFormat.setOutputPath(job, new Path(args[1])); System.exit(job.waitForCompletion(true) ? 0 : 1); } }
3. Python Alternative (Hadoop Streaming)
If you prefer Python, here’s how to implement the same logic with Hadoop Streaming:
Mapper Script
import sys import re for line in sys.stdin: line = line.strip() # Extract all alphabetic words words = re.findall(r'[a-zA-Z]+', line) for word in words: print(f"{len(word)}\t1")
Reducer Script
import sys current_length = None current_count = 0 for line in sys.stdin: line = line.strip() length, count = line.split('\t', 1) try: count = int(count) length = int(length) except ValueError: continue # Skip invalid lines if current_length == length: current_count += count else: if current_length is not None: print(f"{current_length}\t{current_count}") current_length = length current_count = count # Print the last length count if current_length is not None: print(f"{current_length}\t{current_count}")
4. Get Your Expected JSON Output
After running the MapReduce job, you’ll get a text file with lines like:
1 1 2 1 4 3 5 2 9 1
You can easily convert this into your desired JSON format with a simple script (e.g., in Python) that reads the output and constructs the array of objects.
内容的提问来源于stack exchange,提问作者BarZ

