Spring Boot整合MongoDB:MongoRepository无法匹配无重音词汇问题
Hey there! Let's break down why your findByTitleLikeIgnoreCase method isn't returning results when searching for the non-accented "promocao" against a document with title "Christmas Promotion", and walk through practical fixes.
Why This Happens
Spring Data's LikeIgnoreCase translates to a MongoDB query that uses case-insensitive regex matching ($options: 'i'), but this only handles case differences—it doesn't account for accented characters. MongoDB's default string comparison treats accented characters (like ç) and their non-accented counterparts (like c) as distinct values because their underlying Unicode representations are different. That's why searching with the exact accented "promoção" works, but the non-accented version doesn't.
Solutions
1. Use MongoDB Collation for Accent-Insensitive Queries
MongoDB's collation feature lets you define language-specific comparison rules, including ignoring accents. You can add this directly to your repository method with a custom @Query annotation:
@Repository public interface YourDocumentRepository extends MongoRepository<YourDocument, String> { // Using Portuguese locale (matches your example's "promoção") and secondary strength (ignores accents + case) @Query(value = "{ 'title': { $regex: ?0, $options: 'i' } }", collation = "{ 'locale': 'pt', 'strength': 2 }") List<YourDocument> findByTitleAccentInsensitive(String title); }
What the collation settings mean:
locale: 'pt': Optimizes comparison for Portuguese (adjust this to your target language, e.g., 'es' for Spanish, 'fr' for French).strength: 2: Sets comparison to "secondary" level, which ignores both accents and case differences.
Alternatively, if you prefer using MongoTemplate for more control:
@Autowired private MongoTemplate mongoTemplate; public List<YourDocument> searchTitleWithNoAccents(String title) { // Define collation to ignore accents and case Collation collation = Collation.of("pt") .strength(Collation.ComparisonLevel.secondary()); Query query = new Query(Criteria.where("title").regex(title, "i")) .collation(collation); return mongoTemplate.find(query, YourDocument.class); }
2. Preprocess Data with a Normalized Field
Another straightforward approach is to store a normalized version of your title (without accents, lowercase) alongside the original value. This works well if you don't need full-text search capabilities.
First, add a normalized field to your document:
@Document(collection = "your_collection") public class YourDocument { private String title; private String titleNormalized; // Stores title without accents, lowercase // Auto-normalize before saving/updating @PrePersist @PreUpdate private void normalizeTitle() { // Use Apache Commons Lang's stripAccents to remove diacritics this.titleNormalized = StringUtils.stripAccents(this.title).toLowerCase(); } // Getters and setters }
Then add a repository method to query the normalized field:
List<YourDocument> findByTitleNormalizedLike(String normalizedTitle);
When querying, preprocess the input string the same way:
String userInput = "promocao"; String normalizedInput = StringUtils.stripAccents(userInput).toLowerCase(); List<YourDocument> results = repo.findByTitleNormalizedLike("%" + normalizedInput + "%");
3. Use Text Indexes with Collation
If you need full-text search capabilities, create a text index with collation to ignore accents:
@Document(collection = "your_collection") @TextIndexDefinition(fields = { @TextIndexed(field = "title", weight = 1) }) public class YourDocument { private String title; // Other fields }
Then perform a text search with collation:
@Repository public interface YourDocumentRepository extends MongoRepository<YourDocument, String> { @Query(value = "{ $text: { $search: ?0 } }", collation = "{ 'locale': 'pt', 'strength': 2 }") List<YourDocument> findByFullTextAccentInsensitive(String searchTerm); }
内容的提问来源于stack exchange,提问作者Adriano Gomes

