System Design: How to Build an Image Loading Library Like Glide or Coil
Modern Android applications display thousands of images: avatars, product photos, news thumbnails, banners, videos previews, and remote content.
Libraries such as Glide and Coil hide a surprisingly complex system behind a simple API:
ImageLoader(context)
.load("https://example.com/image.jpg")
.into(imageView)
But what happens after load()?
The library may need to:
- Check memory cache.
- Check disk cache.
- Download the image.
- Deduplicate identical requests.
- Decode compressed bytes into a bitmap.
- Resize the image for the target view.
- Apply transformations.
- Avoid blocking the main thread.
- Cancel work when the UI disappears.
- Prevent an old request from replacing a newer one.
- Control memory usage.
- Retry transient network failures.
- Deliver the result safely to the UI.
This article designs such a system from scratch.
The goal is not to reproduce Glide or Coil internally. The goal is to understand the architecture and engineering decisions behind a production-grade Android image-loading library.
1. Requirements
Functional requirements
Our library should support:
- Load images from URLs.
- Load local files/resources.
- Memory caching.
- Disk caching.
- HTTP networking.
- Image decoding.
- Resizing.
- Basic transformations.
- Placeholders and error images.
- Request cancellation.
- Request deduplication.
- Retry support.
- Prefetching.
- Jetpack Compose integration.
- Lifecycle-aware loading.
Non-functional requirements
The library should be:
- Fast
- Memory efficient
- Thread safe
- Lifecycle aware
- Offline friendly
- Extensible
- Testable
- Suitable for large image lists
- Safe against race conditions
2. High-Level Architecture
The complete image-loading pipeline can be represented as:
flowchart LR
UI["Android UI / Compose"] --> REQ["Image Request"]
REQ --> MANAGER["Request Manager"]
MANAGER --> MEM["Memory Cache<br/>LRU"]
MEM -->|Hit| RESULT["Result"]
MEM -->|Miss| DISK["Disk Cache"]
DISK -->|Hit| DECODE["Decoder"]
DISK -->|Miss| NETWORK["Network Fetcher"]
NETWORK --> STORE["Disk Cache Write"]
STORE --> DECODE
DECODE --> RESIZE["Resize / Transform"]
RESIZE --> RESULT
RESULT --> UI
The most important principle is:
Always try the cheapest source first.
The usual lookup order is:
Memory
↓
Disk
↓
Network
↓
Decode
↓
Transform
↓
UI
3. Main Components
A clean architecture could contain these components:
ImageLoader
│
├── RequestManager
│
├── MemoryCache
│
├── DiskCache
│
├── NetworkFetcher
│
├── Decoder
│
├── Transformer
│
├── RequestCoordinator
│
├── Dispatcher
│
└── LifecycleObserver
ImageLoader
Public entry point.
class ImageLoader(
private val memoryCache: MemoryCache,
private val diskCache: DiskCache,
private val networkFetcher: NetworkFetcher,
private val decoder: ImageDecoder
)
The application should not need to know how the internal pipeline works.
4. ImageRequest
Every image operation should be represented by an immutable request.
data class ImageRequest(
val data: String,
val width: Int? = null,
val height: Int? = null,
val transformations: List<Transformation> = emptyList(),
val placeholder: Int? = null,
val error: Int? = null,
val cachePolicy: CachePolicy = CachePolicy.ALL
)
For example:
val request = ImageRequest(
data = "https://example.com/avatar.jpg",
width = 200,
height = 200
)
The request becomes the input to the entire pipeline.
5. Cache Key Design
Caching is not simply:
URL → Bitmap
Because the same URL may be requested with different dimensions or transformations.
For example:
avatar.jpg
avatar.jpg + 100x100
avatar.jpg + 300x300
avatar.jpg + CircleCrop
These can produce different outputs.
A cache key should therefore include the relevant request properties.
data class CacheKey(
val source: String,
val width: Int?,
val height: Int?,
val transformations: List<String>
)
Conceptually:
SHA-256(
source +
width +
height +
transformations
)
This prevents incorrect cache reuse.
6. Memory Cache
Memory cache is the fastest cache.
The most common strategy is an LRU cache.
LRU means:
Least Recently Used
When the cache reaches its maximum size, the least recently accessed item is removed.
Example:
Cache capacity = 4
A B C D
Access A
B C D A
Add E
C D A E
B is removed because it was least recently used.
7. Why Memory Cache Is Important
Without memory caching:
RecyclerView
↓
Download
↓
Decode
↓
Display
Scrolling can trigger expensive operations repeatedly.
With memory cache:
RecyclerView
↓
Memory Cache
↓
Bitmap
This makes repeated image access much faster.
8. Bitmap Memory Is Different From File Size
A common mistake is calculating cache size using the compressed JPEG size.
Suppose:
JPEG = 500 KB
After decoding:
4000 × 3000 × 4 bytes
Approximately:
48 MB
So a 500 KB network file can consume tens of megabytes in memory.
That is why an image-loading library must manage decoded bitmap memory carefully.
9. Simplified MemoryCache
interface MemoryCache {
fun get(key: String): Bitmap?
fun put(
key: String,
bitmap: Bitmap
)
fun remove(key: String)
fun clear()
}
A simplified implementation could use Android's LruCache.
class BitmapMemoryCache(
maxSizeKb: Int
) : MemoryCache {
private val cache = object : LruCache<String, Bitmap>(maxSizeKb) {
override fun sizeOf(
key: String,
bitmap: Bitmap
): Int {
return bitmap.allocationByteCount / 1024
}
}
override fun get(key: String): Bitmap? {
return cache.get(key)
}
override fun put(
key: String,
bitmap: Bitmap
) {
cache.put(key, bitmap)
}
override fun remove(key: String) {
cache.remove(key)
}
override fun clear() {
cache.evictAll()
}
}
The important detail is:
bitmap.allocationByteCount
rather than the compressed file size.
10. Disk Cache
Memory cache disappears when the process is killed.
Disk cache survives process restarts.
A typical architecture is:
Memory Cache
↓ miss
Disk Cache
↓ miss
Network
Disk cache can store:
URL → compressed image bytes
or a processed representation depending on the library design.
For a production implementation, use a robust cache implementation rather than inventing file locking, eviction, and corruption recovery from scratch.
11. Network Layer
The network layer should be isolated behind an interface.
interface NetworkFetcher {
suspend fun fetch(
url: String
): ByteArray
}
An HTTP client such as OkHttp can implement it.
class OkHttpFetcher(
private val client: OkHttpClient
) : NetworkFetcher {
override suspend fun fetch(
url: String
): ByteArray {
val request = Request.Builder()
.url(url)
.build()
client.newCall(request).execute().use { response ->
if (!response.isSuccessful) {
throw IOException(
"HTTP ${response.code}"
)
}
return response.body.bytes()
}
}
}
This separation makes testing easier.
12. Image Decoding
Downloading an image gives us compressed bytes.
The UI needs a decoded representation.
JPEG / PNG / WebP
↓
Image Decoder
↓
Bitmap
A simplified interface:
interface ImageDecoder {
fun decode(
bytes: ByteArray,
width: Int?,
height: Int?
): Bitmap
}
13. Why Downsampling Matters
Imagine a server returns:
4000 × 3000
But the UI displays:
200 × 150
Decoding the full image is wasteful.
The library should ideally decode near the required dimensions.
Network image
↓
Calculate sample size
↓
Decode smaller bitmap
↓
Resize / transform
This reduces:
- Memory usage
- CPU work
- Garbage collection
- Rendering overhead
14. Transformations
Image loading and image transformation are separate responsibilities.
Examples:
CenterCrop
FitCenter
CircleCrop
RoundedCorners
Blur
Grayscale
Interface:
interface Transformation {
fun key(): String
fun transform(
bitmap: Bitmap
): Bitmap
}
Example:
class CircleCrop : Transformation {
override fun key(): String {
return "circle_crop"
}
override fun transform(
bitmap: Bitmap
): Bitmap {
// Transformation implementation
return bitmap
}
}
The transformation key must be part of the cache key.
15. Request Deduplication
This is one of the most important parts of the design.
Imagine a RecyclerView displays 50 items and several items request the same image simultaneously.
Without deduplication:
Request A ──→ Network
Request B ──→ Network
Request C ──→ Network
Three network requests may happen for the same resource.
With request deduplication:
Request A ─┐
Request B ─┼──→ Single Network Request
Request C ─┘
Then the result is shared.
16. In-Flight Request Map
Maintain a map:
private val inFlight =
ConcurrentHashMap<String, Deferred<Bitmap>>()
Conceptually:
CacheKey
↓
Deferred<Bitmap>
When a request arrives:
val existing = inFlight[key]
if (existing != null) {
return existing.await()
}
Otherwise:
val deferred = scope.async {
loadImage(request)
}
inFlight[key] = deferred
try {
return deferred.await()
} finally {
inFlight.remove(key)
}
This can dramatically reduce duplicate work.
17. Complete Loading Pipeline
The central algorithm becomes:
suspend fun load(
request: ImageRequest
): Bitmap {
val key = createCacheKey(request)
memoryCache.get(key)?.let {
return it
}
diskCache.get(key)?.let { bytes ->
val bitmap = decode(request, bytes)
memoryCache.put(key, bitmap)
return bitmap
}
val bytes = networkFetcher.fetch(
request.data
)
diskCache.put(key, bytes)
val bitmap = decode(
request,
bytes
)
val transformed =
applyTransformations(
bitmap,
request.transformations
)
memoryCache.put(
key,
transformed
)
return transformed
}
The production implementation needs additional concerns such as cancellation, synchronization, failures, metrics, and lifecycle handling.
18. Coroutine Architecture
Never perform network or expensive decoding on the main thread.
A simple dispatcher model:
Main
│
├── UI updates
│
IO
│
├── Network
├── Disk
│
Default
│
├── Decode
├── Resize
└── Transform
Example:
withContext(Dispatchers.IO) {
networkFetcher.fetch(url)
}
Then:
withContext(Dispatchers.Default) {
decoder.decode(bytes)
}
The exact dispatcher strategy should be tuned based on the underlying decoder and workload.
19. Structured Concurrency
A good image loader should respect coroutine cancellation.
For example:
val job = scope.launch {
imageLoader.load(request)
}
If the UI no longer needs the image:
job.cancel()
The loader should allow cancellation to propagate to:
Coroutine
↓
Network
↓
Decode
↓
Transformation
This avoids unnecessary work.
20. Lifecycle Awareness
Suppose an Activity starts loading:
Image A
Image B
Image C
The user navigates away.
Continuing all work may waste:
- Network bandwidth
- CPU
- Memory
- Battery
The UI layer should therefore associate requests with lifecycle-aware scopes.
In Compose:
LaunchedEffect(model) {
imageLoader.load(
ImageRequest(model)
)
}
When the effect leaves composition, the coroutine can be cancelled.
21. The RecyclerView Problem
A classic problem:
Position 0 → image A
Position 1 → image B
Position 2 → image C
The view is reused:
Position 0 → image D
But request A finishes late.
Without protection:
Old Request A
↓
Reused View
↓
Wrong image displayed
The request must be associated with the current target.
Conceptually:
ImageView
↓
Request Token
Before delivering a result:
if (target.currentRequestId == requestId) {
target.setImageBitmap(bitmap)
}
This prevents stale results.
22. Error Handling
Network requests can fail.
Possible failures:
DNS failure
Timeout
HTTP 404
HTTP 500
Connection reset
Invalid image
OutOfMemoryError
Cancellation
Disk failure
Use a structured result:
sealed interface ImageResult {
data class Success(
val bitmap: Bitmap
) : ImageResult
data class Error(
val throwable: Throwable
) : ImageResult
}
The UI can then decide what to display.
23. Retry Strategy
Not every failure should be retried.
For example:
404 → usually don't retry
500 → may retry
Timeout → may retry
Cancellation → don't retry
Invalid image → don't retry
A simple exponential backoff:
Attempt 1 → 250 ms
Attempt 2 → 500 ms
Attempt 3 → 1000 ms
Add jitter in a real implementation to avoid synchronized retry storms.
24. Prefetching
Suppose a feed is currently displaying:
Item 1
Item 2
Item 3
The user will probably scroll toward:
Item 4
Item 5
Item 6
The library can prefetch them.
Visible items
↓
Predict upcoming items
↓
Prefetch
↓
Disk / Memory Cache
When the user reaches item 4:
Memory Cache
↓
Instant result
Prefetching should remain bounded so it does not consume excessive bandwidth or memory.
25. Jetpack Compose API
A modern library should expose a Compose-friendly API.
For example:
@Composable
fun AsyncImage(
model: String,
contentDescription: String?,
modifier: Modifier = Modifier
)
Usage:
AsyncImage(
model = "https://example.com/avatar.jpg",
contentDescription = "Avatar"
)
Internally:
Composable
↓
ImageRequest
↓
ImageLoader
↓
Cache / Network
↓
Bitmap
↓
State
↓
Composable
26. Compose State
A simplified implementation:
@Composable
fun SimpleAsyncImage(
url: String,
imageLoader: ImageLoader
) {
var bitmap by remember(url) {
mutableStateOf<Bitmap?>(null)
}
LaunchedEffect(url) {
bitmap = imageLoader.load(
ImageRequest(url)
)
}
bitmap?.let {
Image(
bitmap = it.asImageBitmap(),
contentDescription = null
)
}
}
A production library needs placeholders, errors, cancellation, state transitions, content scaling, accessibility, and lifecycle handling.
27. Threading Model
A useful mental model:
ImageLoader
│
┌──────────┴──────────┐
│ │
Memory Request
Cache Coordinator
│
┌────────┼────────┐
│ │ │
Disk Network Decode
IO IO CPU
│ │ │
└────────┼────────┘
│
Transformation
│
UI / Compose
The key idea is to separate:
- IO-bound work
- CPU-bound work
- UI work
28. Complete System Design
flowchart TB
UI["Jetpack Compose / Views"]
UI --> API["ImageLoader API"]
API --> RM["Request Manager"]
RM --> KEY["Cache Key Generator"]
KEY --> MEM["Memory Cache<br/>LRU"]
MEM -->|Hit| RESULT["Result"]
MEM -->|Miss| DEDUP["In-Flight Request Deduplication"]
DEDUP --> DISK["Disk Cache"]
DISK -->|Hit| DECODER["Image Decoder"]
DISK -->|Miss| HTTP["HTTP Client"]
HTTP --> DISK
HTTP -->|Bytes| DECODER
DECODER --> RESIZE["Downsampling / Resize"]
RESIZE --> TRANSFORM["Transformations"]
TRANSFORM --> MEM
TRANSFORM --> RESULT
RESULT --> UI
RM --> LIFE["Lifecycle / Cancellation"]
RM --> RETRY["Retry / Backoff"]
29. Suggested Package Structure
A clean Android library could use:
image-loader/
│
├── core/
│ ├── ImageLoader.kt
│ ├── ImageRequest.kt
│ ├── ImageResult.kt
│ └── CacheKey.kt
│
├── cache/
│ ├── MemoryCache.kt
│ ├── BitmapMemoryCache.kt
│ └── DiskCache.kt
│
├── network/
│ ├── NetworkFetcher.kt
│ └── OkHttpFetcher.kt
│
├── decode/
│ ├── ImageDecoder.kt
│ └── BitmapDecoder.kt
│
├── transform/
│ ├── Transformation.kt
│ ├── CenterCrop.kt
│ └── CircleCrop.kt
│
├── request/
│ ├── RequestManager.kt
│ └── RequestCoordinator.kt
│
├── compose/
│ └── AsyncImage.kt
│
└── lifecycle/
└── LifecycleObserver.kt
This separation keeps the core engine independent from UI integrations.
30. Important System Design Trade-offs
Memory vs Performance
More memory cache:
+ Faster image access
- Higher memory usage
Less memory cache:
+ Lower memory pressure
- More disk/network work
Disk vs Network
More disk caching:
+ Less network usage
+ Better offline behavior
- More storage
Less disk caching:
+ Less storage
- More network traffic
Full Resolution vs Downsampling
Full-resolution decoding:
+ Maximum detail
- High memory consumption
- Higher CPU cost
Downsampling:
+ Lower memory
+ Faster decoding
- May reduce available detail
31. What Happens During One Image Request?
Suppose:
URL:
https://example.com/cat.jpg
Target:
300 × 300
Transformation:
CircleCrop
The request travels through:
1. Create ImageRequest
2. Generate cache key
3. Check Memory Cache
4. If miss → Check Disk Cache
5. If miss → Fetch network bytes
6. Store bytes in Disk Cache
7. Calculate decode size
8. Decode bitmap
9. Resize
10. Apply CircleCrop
11. Store final result in Memory Cache
12. Deliver result to UI
This is the core architecture behind a modern image loader.
32. How to Make It Production Ready
A production-grade implementation should additionally consider:
Memory pressure
Respond to Android memory pressure events.
onTrimMemory()
↓
Reduce / clear memory cache
HTTP caching
Respect:
Cache-Control
ETag
Last-Modified
Expires
This can avoid unnecessary downloads.
Metrics
Track:
Memory cache hit rate
Disk cache hit rate
Network requests
Decode duration
Transformation duration
Average image size
Failures
Cancellation rate
Observability
Expose optional debug logging:
[ImageLoader]
CACHE_MEMORY_HIT
CACHE_DISK_MISS
NETWORK_START
NETWORK_SUCCESS
DECODE_START
DECODE_SUCCESS
TRANSFORM
DELIVER
This makes performance problems much easier to diagnose.
33. Scaling Considerations
For normal Android applications, the architecture above is sufficient.
For a very large application, additional optimizations can include:
Request prioritization
↓
Bounded concurrency
↓
Connection pooling
↓
HTTP/2 / HTTP/3
↓
CDN
↓
Image resizing at the server
↓
WebP / AVIF
↓
Progressive loading
Server-side image resizing is particularly valuable.
Instead of:
Server → 4000 × 3000
request:
Server → 300 × 225
when the UI only needs a thumbnail.
34. Common System Design Interview Questions
Why do we need both memory and disk cache?
Memory is much faster but limited and process-bound. Disk is slower but persistent across process restarts.
Why use LRU?
Because recently used images are more likely to be requested again, while the cache needs bounded memory.
Why is request deduplication important?
Multiple UI components may request the same resource concurrently. Deduplication avoids duplicate network, decoding, and transformation work.
Why include transformations in the cache key?
The same source image can produce different output images.
Why downsample?
To avoid decoding huge images when the UI only needs a small bitmap.
Why lifecycle-aware cancellation?
Because UI components can disappear before a request completes.
Why separate decoding from networking?
They have different resource characteristics and can evolve independently.
35. Learning Path
If you want to implement this library yourself, build it incrementally.
Phase 1 — Basic loader
Implement:
URL
↓
HTTP
↓
Bitmap
↓
ImageView
Phase 2 — Memory cache
Add:
LruCache
Phase 3 — Disk cache
Add:
DiskCache
Phase 4 — Coroutines
Introduce:
suspend
Dispatchers.IO
Dispatchers.Default
Cancellation
Phase 5 — Request deduplication
Implement:
ConcurrentHashMap
Deferred
Phase 6 — Image transformations
Implement:
Resize
CenterCrop
CircleCrop
RoundedCorners
Phase 7 — Compose
Build:
AsyncImage(...)
Phase 8 — Production features
Add:
Retry
Prefetch
Lifecycle
Metrics
Memory pressure
HTTP cache validation
Request priority
36. Final Architecture
The final mental model is:
┌───────────────────┐
│ Compose / View │
└─────────┬─────────┘
│
ImageRequest
│
┌─────────▼─────────┐
│ RequestManager │
└─────────┬─────────┘
│
Generate Key
│
┌────────────▼────────────┐
│ Memory Cache │
└────────────┬────────────┘
│ miss
┌────────────▼────────────┐
│ Disk Cache │
└────────────┬────────────┘
│ miss
┌────────────▼────────────┐
│ Network Layer │
└────────────┬────────────┘
│
┌────────────▼────────────┐
│ Image Decoder │
└────────────┬────────────┘
│
┌────────────▼────────────┐
│ Resize / Transformation │
└────────────┬────────────┘
│
┌────────────▼────────────┐
│ Memory Cache │
└────────────┬────────────┘
│
Result → UI
The key lesson is that an image-loading library is not simply a network utility.
It is a concurrent resource-management system combining:
- Caching
- Networking
- Image decoding
- Memory management
- Concurrency
- Cancellation
- Lifecycle awareness
- Transformations
- UI state management
Once you understand this pipeline, libraries like Glide and Coil become much easier to reason about—not because their implementations are simple, but because their responsibilities become clear.
Conclusion
Designing an image-loading library is an excellent Android system-design exercise because it connects low-level Android concepts with large-scale architecture.
The most important pieces to understand are:
Memory Cache
+
Disk Cache
+
Network
+
Image Decoding
+
Request Deduplication
+
Concurrency
+
Lifecycle
+
Transformations
=
Production Image Loader
Start with a minimal loader, then add one capability at a time. This approach teaches far more than simply using an existing library.
Author
Padmakar Garg
Android Developer focused on Kotlin, Jetpack Compose, Android Architecture, System Design, and Kotlin Multiplatform.
GitHub: https://github.com/gargpadmakar
Buy Me a Coffee: https://buymeacoffee.com/padmakargarg
Top comments (0)