You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DynamoDB中存储翻译并最小化重复、优化查询效率?

Is Your Current DynamoDB Schema Unreasonable? Better Alternatives for Storing Translations

Yep, your current DynamoDB schema is definitely not a good fit—let’s break down why, then walk through some way better approaches tailored to your core query need: finding all objects that contain a specific translation value.

Why Your Current Setup Falls Short

  • Massive Wasted Storage: Storing 6500 language entries (99% of which are empty) bloats your items for no reason. DynamoDB caps each item at 400KB, so you’ll hit that limit fast with this structure, even for simple fruit entries.
  • Slow, Inefficient Queries: To find a specific translation, you’d have to scan or query against an array of single-key objects. This forces DynamoDB to iterate through every element in the array for each item—wasting read capacity units (RCUs) and slowing down responses.
  • Clunky Query Logic: Checking for a translation like "apelsin" would require awkward conditions like contains(fruit.translations, {"sv-SE": "apelsin"}), which is error-prone and doesn’t scale as you add more languages.

Optimized Schema Designs

1. Flattened Translation Map (Great for Direct Language Lookups)

Ditch the array of empty objects and store translations as a nested map that only includes languages with actual content:

{
  "timestamp": "2024-05-20T12:00:00Z",
  "fruit": {
    "name": "orange",
    "translations": {
      "en-GB": "orange",
      "sv-SE": "apelsin"
      // Only add languages with translations—no empty entries
    }
  }
}
  • Pros: Saves tons of space, makes it super easy to look up a translation by language code (e.g., fruit.translations.sv-SE), and keeps items compact.
  • Cons: Not great for your reverse query (finding all items with a specific translation value), since DynamoDB can’t efficiently search across all map values without a full scan.

2. Split into a Dedicated Translations Table (Best for Reverse Queries)

DynamoDB works best when you model tables around your access patterns. For your need to find items by translation content, split your data into two tables:

Main Fruit Table

Stores core fruit metadata:

{
  "fruit_name": "orange", // Partition key
  "timestamp": "2024-05-20T12:00:00Z"
  // Add other fruit details here if needed
}

Translations Table

Stores one entry per fruit-language pair:

{
  "fruit_name": "orange", // Partition key
  "lang_code": "sv-SE",   // Sort key
  "translation": "apelsin",
  "timestamp": "2024-05-20T12:00:00Z"
}

Then add a Global Secondary Index (GSI) to the Translations Table with:

  • Partition key: translation

  • Projection: Include fruit_name (and any other fields you need to retrieve directly)

  • How to Query: To find all items with the translation "apelsin", run a point query against the GSI’s partition key for "apelsin". You’ll get all matching fruit_name values, then fetch the full fruit details from the main table if required.

  • Pros: Perfectly optimized for your reverse query need, scales infinitely with new languages, no wasted space, and uses efficient point queries that are easy on RCUs.

  • Cons: Requires two table operations if you need full fruit details, but this is a tiny tradeoff for the performance and scalability gains.

3. Hybrid: Combined Item with Translation Values List

If you prefer keeping all data in one item, add a list of just the translation values and index it with a GSI:

{
  "timestamp": "2024-05-20T12:00:00Z",
  "fruit": {
    "name": "orange",
    "translations": {
      "en-GB": "orange",
      "sv-SE": "apelsin"
    }
  },
  "translation_values": ["orange", "apelsin"] // Sync this list with translations
}

Create a GSI where translation_values is the partition key (or sort key, depending on your other query needs).

  • How to Query: Use contains(translation_values, "apelsin") in a query against the GSI to find matching items.
  • Pros: Keeps all data in a single item, supports reverse queries with a GSI.
  • Cons: You have to manually keep the translation_values list in sync with the translations map, and the list can get large if a fruit has many translations (though still way better than storing empty entries).

Final Recommendation

If your top priority is efficiently finding items by their translation content, go with the dedicated translations table + GSI—it’s the most scalable, efficient, and aligned with DynamoDB’s best practices. If you mostly look up translations by language code and rarely need reverse queries, the flattened map is simpler and more straightforward.

内容的提问来源于stack exchange,提问作者Amc_rtty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:43:59