Apache Solr with Java – Building High-Performance Search Solutions
Introduction
In modern distributed systems and microservices architectures, search functionality is often a critical requirement. Whether you're building e-commerce platforms, financial systems, or content management solutions, your application needs to handle complex queries efficiently across large datasets. Apache Solr, built on top of Apache Lucene, provides exactly that—a powerful, open-source search platform that integrates seamlessly with Java applications.
This comprehensive guide walks you through everything you need to know about integrating Apache Solr with Java, from basic setup to production-grade implementations. Whether you're a fintech architect designing search solutions for banking systems or a backend engineer optimizing search performance in microservices, this guide covers real-world scenarios and best practices.
What is Apache Solr?
Apache Solr is a fast, open-source search platform built on top of Apache Lucene. It provides:
- Full-text search capabilities with advanced relevance scoring
- Distributed search across multiple nodes and shards
- Real-time indexing with low-latency queries
- Faceted search for filtering and navigation
- Advanced query syntax supporting complex boolean, wildcard, and range queries
- REST API for language-agnostic integration
- High availability through replication and clustering
Unlike embedded search solutions, Solr runs as a standalone service, allowing multiple applications to share a centralized search infrastructure. This is particularly valuable in enterprise environments where search consistency and scalability are paramount.
Why Solr for Java Applications?
Performance at Scale
Solr handles billions of documents efficiently. In fintech systems processing trading data, market feeds, and transaction histories, Solr can index and query millions of records per second while maintaining sub-100ms response times.
Java Integration
Solr provides official Java clients (SolrJ) that offer:
- Type-safe query construction
- Connection pooling and failover handling
- Automatic retry logic
- Efficient serialization using binary protocols
Enterprise Features
- Secure authentication and authorization via security plugins
- Request rate limiting and circuit breakers
- Monitoring and metrics integration with common tools
- Transaction logs for durability and recovery
Cost-Effectiveness
As an open-source solution, Solr eliminates licensing costs while maintaining enterprise-grade reliability.
Setting Up Apache Solr
Installation
# Download Solr (as of 2024, latest stable is 9.x)
wget https://archive.apache.org/dist/solr/solr-9.3.0/solr-9.3.0.tgz
tar xzf solr-9.3.0.tgz
cd solr-9.3.0
# Start Solr in standalone mode (development)
bin/solr start
# Or start in SolrCloud mode (production) with Zookeeper
bin/solr start -c -z localhost:2181
Creating a Collection
# Create a collection with name 'products' and 2 shards
bin/solr create -c products -s 2 -rf 2
# Verify collection
curl http://localhost:8983/solr/products/admin/ping
Java Client Integration: SolrJ
Maven Dependency
<dependency>
<groupId>org.apache.solr</groupId>
<artifactId>solr-solrj</artifactId>
<version>9.3.0</version>
</dependency>
Basic Configuration
import org.apache.solr.client.solrj.SolrClient;
import org.apache.solr.client.solrj.impl.HttpSolrClient;
import org.apache.solr.client.solrj.impl.CloudSolrClient;
public class SolrClientConfiguration {
// Standalone Solr client
public static SolrClient createHttpClient(String solrUrl) {
return new HttpSolrClient.Builder(solrUrl)
.withConnectionTimeout(10000)
.withSocketTimeout(60000)
.build();
}
// SolrCloud client (distributed)
public static SolrClient createCloudClient(String zkHosts) {
return new CloudSolrClient.Builder()
.withZkHost(zkHosts)
.build();
}
}
Indexing Documents
Document Model
import org.apache.solr.client.solrj.beans.Field;
public class Product {
@Field("id")
private String id;
@Field("name")
private String name;
@Field("description")
private String description;
@Field("price")
private Double price;
@Field("category")
private String category;
@Field("in_stock")
private Boolean inStock;
@Field("timestamp")
private Long timestamp;
// Constructor, getters, setters...
public Product() {}
public Product(String id, String name, String description, Double price) {
this.id = id;
this.name = name;
this.description = description;
this.price = price;
this.timestamp = System.currentTimeMillis();
}
}
Indexing with Batch Processing
import org.apache.solr.client.solrj.SolrClient;
import org.apache.solr.client.solrj.response.UpdateResponse;
public class SolrIndexingService {
private final SolrClient solrClient;
private final String collectionName = "products";
private static final int BATCH_SIZE = 1000;
public SolrIndexingService(SolrClient solrClient) {
this.solrClient = solrClient;
}
/**
* Index products in batches for optimal performance
*/
public void indexProducts(List<Product> products) throws Exception {
for (int i = 0; i < products.size(); i += BATCH_SIZE) {
int end = Math.min(i + BATCH_SIZE, products.size());
List<Product> batch = products.subList(i, end);
UpdateResponse response = solrClient.addBeans(
collectionName,
batch,
5000 // commit within 5 seconds
);
if (response.getStatus() != 0) {
throw new RuntimeException("Indexing failed: " + response.getResponseHeader());
}
System.out.println("Indexed batch: " + (i + batch.size()) + "/" + products.size());
}
// Final commit
solrClient.commit(collectionName);
}
/**
* Delete products by ID
*/
public void deleteProduct(String id) throws Exception {
solrClient.deleteById(collectionName, id);
solrClient.commit(collectionName);
}
/**
* Delete all products matching a query
*/
public void deleteByQuery(String query) throws Exception {
solrClient.deleteByQuery(collectionName, query);
solrClient.commit(collectionName);
}
}
Querying Solr
Basic Query Operations
import org.apache.solr.client.solrj.SolrQuery;
import org.apache.solr.client.solrj.response.QueryResponse;
import org.apache.solr.common.SolrDocumentList;
public class SolrSearchService {
private final SolrClient solrClient;
private final String collectionName = "products";
public SolrSearchService(SolrClient solrClient) {
this.solrClient = solrClient;
}
/**
* Full-text search across all fields
*/
public List<Product> search(String queryString, int start, int rows) throws Exception {
SolrQuery query = new SolrQuery();
query.setQuery(queryString);
query.setStart(start);
query.setRows(rows);
query.addFilterQuery("in_stock:true"); // Only in-stock products
QueryResponse response = solrClient.query(collectionName, query);
return response.getBeans(Product.class);
}
/**
* Advanced query with faceting
*/
public SearchResult searchWithFacets(String queryString) throws Exception {
SolrQuery query = new SolrQuery();
query.setQuery(queryString);
// Add facets for filtering
query.addFacetField("category");
query.addFacetField("price");
query.setFacetSort("count");
query.setFacetLimit(20);
QueryResponse response = solrClient.query(collectionName, query);
return new SearchResult(
response.getBeans(Product.class),
response.getFacetFields()
);
}
/**
* Range query with boost
*/
public List<Product> searchByPrice(Double minPrice, Double maxPrice) throws Exception {
SolrQuery query = new SolrQuery();
query.setQuery("*:*"); // Match all
query.addFilterQuery("price:[" + minPrice + " TO " + maxPrice + "]");
query.setSort("price", SolrQuery.ORDER.asc);
query.setRows(100);
QueryResponse response = solrClient.query(collectionName, query);
return response.getBeans(Product.class);
}
}
Advanced Query Syntax
/**
* Examples of advanced Solr query syntax
*/
public class SolrQueryExamples {
// Boolean query
private static final String BOOL_QUERY =
"name:laptop AND (category:electronics OR category:computers)";
// Wildcard query
private static final String WILDCARD_QUERY = "name:solr*";
// Phrase query (exact match)
private static final String PHRASE_QUERY = "description:\"high performance search\"";
// Proximity query (words within distance)
private static final String PROXIMITY_QUERY = "description:\"search engine\"~5";
// Range query with date
private static final String DATE_RANGE =
"timestamp:[2024-01-01T00:00:00Z TO 2024-12-31T23:59:59Z]";
// Fuzzy search (typo tolerance)
private static final String FUZZY_QUERY = "name:solr~1";
// Boost specific fields
private static final String BOOST_QUERY =
"name:solr^2.0 OR description:solr^1.0";
}
Real-World Production Example: Banking Search
Here's a practical example from fintech systems—searching transactions:
import org.apache.solr.client.solrj.SolrQuery;
import java.time.LocalDateTime;
import java.time.format.DateTimeFormatter;
public class TransactionSearchService {
private final SolrClient solrClient;
private final String collectionName = "transactions";
private static final DateTimeFormatter ISO_FORMATTER =
DateTimeFormatter.ISO_DATE_TIME;
/**
* Search transactions within date range with amount filter
*/
public List<Transaction> searchTransactions(
String accountId,
LocalDateTime startDate,
LocalDateTime endDate,
Double minAmount,
Double maxAmount) throws Exception {
SolrQuery query = new SolrQuery();
query.setQuery("account_id:" + accountId);
// Add filters for date and amount
query.addFilterQuery(
"timestamp:[" + startDate.format(ISO_FORMATTER) +
" TO " + endDate.format(ISO_FORMATTER) + "]"
);
query.addFilterQuery(
"amount:[" + minAmount + " TO " + maxAmount + "]"
);
query.setRows(1000);
query.setSort("timestamp", SolrQuery.ORDER.desc);
QueryResponse response = solrClient.query(collectionName, query);
return response.getBeans(Transaction.class);
}
/**
* Full-text search on transaction descriptions
*/
public List<Transaction> searchDescriptions(
String accountId,
String searchTerm) throws Exception {
SolrQuery query = new SolrQuery();
query.setQuery("account_id:" + accountId +
" AND description:" + searchTerm);
// Highlight matching terms
query.setHighlight(true);
query.addHighlightField("description");
query.setHighlightFragsize(150);
QueryResponse response = solrClient.query(collectionName, query);
return response.getBeans(Transaction.class);
}
}
@Field("account_id")
public class Transaction {
@Field("id")
private String id;
@Field("account_id")
private String accountId;
@Field("amount")
private BigDecimal amount;
@Field("description")
private String description;
@Field("timestamp")
private LocalDateTime timestamp;
@Field("status")
private String status;
}
Performance Optimization & Best Practices
1. Connection Pooling
import org.apache.http.impl.client.HttpClientBuilder;
import org.apache.http.impl.conn.PoolingHttpClientConnectionManager;
public class SolrClientFactory {
public static SolrClient createOptimizedClient(String solrUrl) {
PoolingHttpClientConnectionManager cm =
new PoolingHttpClientConnectionManager();
cm.setMaxTotal(100);
cm.setDefaultMaxPerRoute(50);
return new HttpSolrClient.Builder(solrUrl)
.withHttpClient(
HttpClientBuilder.create()
.setConnectionManager(cm)
.build()
)
.withConnectionTimeout(10000)
.withSocketTimeout(60000)
.build();
}
}
2. Batch Indexing Strategy
- Index in batches of 500-1000 documents
- Use commit intervals (e.g., every 5 seconds)
- Implement exponential backoff for retries
- Monitor heap usage and adjust batch size accordingly
3. Query Optimization
// ✅ GOOD: Use filter queries for high-cardinality fields
query.addFilterQuery("status:ACTIVE");
query.addFilterQuery("country:US");
// ✅ GOOD: Specify required fields only
query.setFields("id", "name", "price");
// ❌ BAD: Avoid expensive wildcards at start
// query.setQuery("*laptop*"); // Very slow!
// ✅ GOOD: Use leading wildcard sparingly
query.setQuery("laptop*");
4. Monitoring & Circuit Breaker Pattern
import io.github.resilience4j.circuitbreaker.CircuitBreaker;
import io.github.resilience4j.core.registry.EntryAddedEvent;
public class ResilientSolrClient {
private final CircuitBreaker circuitBreaker;
private final SolrClient solrClient;
public ResilientSolrClient(SolrClient solrClient) {
this.solrClient = solrClient;
this.circuitBreaker = CircuitBreaker.ofDefaults("solr-search");
}
public QueryResponse search(String collection, SolrQuery query)
throws Exception {
return circuitBreaker.executeSupplier(() ->
solrClient.query(collection, query)
);
}
}
Solr vs Elasticsearch: When to Choose Solr
| Feature | Solr | Elasticsearch |
|---|---|---|
| Java Integration | Native SolrJ | Java client available |
| Operational Complexity | Lower | Higher |
| Community Size | Stable | Very Large |
| Out-of-box Features | Rich (faceting, highlighting) | Requires plugins |
| Learning Curve | Moderate | Steep |
| Cloud Deployment | SolrCloud (mature) | Native cloud support |
| Cost | Free/Open-source | Free/Open-source |
Choose Solr when you need:
- Rapid deployment with minimal configuration
- Strong faceted search capabilities
- Direct Java integration with type safety
- Established financial/banking systems
Troubleshooting Common Issues
Issue: Slow Indexing Performance
Solution:
- Increase batch size
- Add more commit intervals
- Check heap memory allocation
- Monitor GC pauses
Issue: High Query Latency
Solution:
- Add filter queries instead of boolean clauses
- Implement query result caching
- Increase JVM heap for filter cache
- Consider adding replicas for read distribution
Issue: OutOfMemory Errors
Solution:
- Reduce batch indexing size from 1000 to 500
- Increase -Xmx JVM parameter
- Implement warming queries before use
- Monitor with JProfiler or YourKit
Conclusion
Apache Solr remains a powerful choice for building high-performance search solutions in Java applications. Whether you're designing search infrastructure for fintech platforms, e-commerce systems, or content management solutions, Solr's maturity, flexibility, and enterprise features make it an excellent option.
The key to successful Solr implementation is:
- Proper architecture design with sharding strategy
- Efficient indexing using batch processing
- Optimized queries with filter queries and field selection
- Robust error handling with circuit breakers
- Continuous monitoring and performance tuning
By following the patterns and best practices outlined in this guide, you'll build search solutions that scale reliably and deliver sub-100ms response times even under heavy load.
Happy searching!
Top comments (0)