MongoDB是否提供SimplifiedJson转带原字段类型的Bson内置方法?
Let's break down how to resolve this issue where your SimplifiedJson output loses critical type information (like UUIDs being parsed as strings instead of BsonBinary):
1. Fix at the Source: Output ExtendedJson Directly
The most straightforward and reliable fix is to adjust your MongoDB Kafka Connector configuration to output ExtendedJson instead of SimplifiedJson. ExtendedJson includes explicit type metadata that the MongoDB Java SDK can natively parse into the correct Bson types without extra client-side work.
Update your connector's output.json.formatter setting to:
output.json.formatter=com.mongodb.kafka.connect.source.json.formatter.ExtendedJson
With this change, your UUID _id will serialize to:
"_id": {"$binary": {"base64": "MSRJCs07SFy4sMpopdRvEA==", "subType": "04"}}
When you run BsonDocument.parse() on this, the SDK will automatically recognize it as a BsonBinary (with UUID subtype 04) instead of a string. This works seamlessly for all Bson types—ObjectId, Date, numeric types, and more—no custom code required.
2. If You Must Use SimplifiedJson: Enable Connector Schemas
If switching to ExtendedJson isn't an option, enable schema support in the Kafka Connector to include type metadata alongside simplified JSON values. This gives your client the context needed to deserialize values back to their original Bson types.
Adjust your connector's converter settings:
value.converter=org.apache.kafka.connect.json.JsonConverter value.converter.schemas.enable=true
With schemas enabled, each Kafka message will include a schema field that defines the type of every value (e.g., binary for UUIDs, objectid for ObjectIds). You can then use this schema to guide deserialization:
- Parse the schema from the incoming message.
- Map each field's schema type to its corresponding Bson type.
- Convert the simplified JSON value to the correct Bson type (e.g., decode the base64 string to
BsonBinaryfor UUIDs).
Why SimplifiedJson Creates This Problem
SimplifiedJson is optimized for readability, so it converts Bson types to plain JSON values without type markers:
- UUIDs become base64-encoded strings
- ObjectIds become hex strings
- Dates become ISO-8601 strings
Since there's no type metadata included, BsonDocument.parse() has no way to infer the original type—it treats all these values as plain strings. The MongoDB Java SDK doesn't have a built-in method to "guess" the correct type from unmarked simplified JSON, as the same string could represent multiple types (e.g., a hex string might be an ObjectId or a regular string).
内容的提问来源于stack exchange,提问作者Marlon Patrick

