自定义Hadoop Writable二次排序时readFields方法抛出EOFException求助
Hey there! That EOFException popping up in your readFields() method is almost always a red flag that your serialization and deserialization logic isn't perfectly in sync. Let's break down the most common causes and how to fix them:
1. Mismatched write() and readFields() Methods
This is by far the #1 culprit. Every piece of data you write in the write() method must be read in exactly the same order, with exactly the same data type, and using matching I/O methods in readFields().
For example, if your write() does this:
public void write(DataOutput out) throws IOException { out.writeInt(primaryKey); out.writeLong(secondaryKey); out.writeUTF(someValue); }
Your readFields() must mirror it perfectly:
public void readFields(DataInput in) throws IOException { primaryKey = in.readInt(); secondaryKey = in.readLong(); someValue = in.readUTF(); }
If you swap the order of readLong() and readInt(), or forget to read one of the fields, Hadoop will try to read beyond the end of the serialized data—boom, EOFException.
2. Improper Handling of Variable-Length Data
When dealing with strings or byte arrays, it's easy to mess up length tracking. Here's what to avoid:
- Don't use
writeBytes()for strings unless you explicitly write the length first—readUTF()expects a prefixed length, whichwriteBytes()doesn't provide. - For custom byte arrays, always write the length before the bytes themselves, then read the length first to know how many bytes to fetch.
Bad example:
// Wrong write public void write(DataOutput out) throws IOException { out.writeBytes(myString); // No length prefix } // Wrong read public void readFields(DataInput in) throws IOException { myString = in.readUTF(); // Tries to read a non-existent length value }
Fixed version:
// Correct write public void write(DataOutput out) throws IOException { out.writeUTF(myString); // Automatically handles length prefixing } // Correct read public void readFields(DataInput in) throws IOException { myString = in.readUTF(); }
Or for raw byte arrays:
public void write(DataOutput out) throws IOException { byte[] data = myData.getBytes(StandardCharsets.UTF_8); out.writeInt(data.length); // Write length first out.write(data); } public void readFields(DataInput in) throws IOException { int length = in.readInt(); // Read length first byte[] data = new byte[length]; in.readFully(data); // Read exact number of bytes myData = new String(data, StandardCharsets.UTF_8); }
3. Unhandled Null Values
If your Writable has fields that can be null, you need to explicitly handle them during serialization. Without a marker for nulls, readFields() will still try to read data that isn't there:
public void write(DataOutput out) throws IOException { out.writeInt(primaryKey); if (secondaryValue == null) { out.writeBoolean(false); } else { out.writeBoolean(true); out.writeUTF(secondaryValue); } } public void readFields(DataInput in) throws IOException { primaryKey = in.readInt(); boolean hasValue = in.readBoolean(); if (hasValue) { secondaryValue = in.readUTF(); } else { secondaryValue = null; } }
4. Test Serialization in Isolation
Don't debug this by running the entire Hadoop job—write a quick unit test to validate your Writable's logic in isolation. This will let you pinpoint issues fast:
import org.junit.Assert; import org.junit.Test; import java.io.*; public class MyCustomWritableTest { @Test public void testSerialization() throws IOException { // Create original object MyCustomWritable original = new MyCustomWritable(100, 200L, "test-data"); // Serialize to byte array ByteArrayOutputStream baos = new ByteArrayOutputStream(); DataOutputStream dos = new DataOutputStream(baos); original.write(dos); // Deserialize from byte array DataInputStream dis = new DataInputStream(new ByteArrayInputStream(baos.toByteArray())); MyCustomWritable deserialized = new MyCustomWritable(); deserialized.readFields(dis); // Verify fields match Assert.assertEquals(original.getPrimaryKey(), deserialized.getPrimaryKey()); Assert.assertEquals(original.getSecondaryKey(), deserialized.getSecondaryKey()); Assert.assertEquals(original.getSomeValue(), deserialized.getSomeValue()); } }
Start with checking the alignment of write() and readFields()—9 times out of 10, that's where the issue lies. The unit test will quickly confirm if your serialization logic is working as expected.
内容的提问来源于stack exchange,提问作者yzhan

